Search NASA⌕ Search

SEARCH · Search NASA

Results for “Conserved Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A novel bacterial protein family that catalyses nitrous oxide reduction

Nitrous oxide (N 2 O), a driver of global warming and climate change, has reached unprecedented concentrations in Earth’s atmosphere. Current N 2 O sources outpace N 2 O sinks, emphasizing the need for comprehensive understanding of processes that consume N 2 O. Microbes that express the enzyme N 2 O reductase (N 2 OR) convert N 2 O to climate change-neutral dinitrogen (N 2 ). Known N 2 ORs belong to the canonical clade I and clade II NosZ reductases and are considered key enzymes for N 2 O reduction. Here we report a previously unrecognized protein family with a role in N 2 O reduction, clade III lactonase-type N 2 OR (L-N 2 OR), which diverges in sequence from canonical NosZ but conserves three-dimensional protein structural features. Integrated physiological, metagenomic, proteomic and structural modelling studies demonstrate that L-N 2 ORs catalyse N 2 O reduction. L-N 2 OR genes occur in several phyla, predominantly in uncultured taxa with broad geographic distribution. Our findings expand the known diversity of N 2 ORs and implicate previously unrecognized taxa (for example, Nitrospinota) in N 2 O consumption. In conclusion, the expansion of N 2 OR diversity and the identification of a novel type of catalyst for N 2 O reduction advances the understanding of N 2 O sinks, has implications for greenhouse gas emission and climate change modelling, and expands opportunities for innovative biotechnologies aimed at curbing N 2 O emissions.

He, Guang 何广 [Univ. of Tennessee, Knoxville, TN (U↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

A ligand discovery toolbox for the WWE domain family of human E3 ligases

The WWE domain is a relatively under-researched domain found in twelve human proteins and characterized by a conserved tryptophan-tryptophan-glutamate (WWE) sequence motif. Six of these WWE domain-containing proteins also contain domains with E3 ubiquitin ligase activity. The general recognition of poly-ADP-ribosylated substrates by WWE domains suggests a potential avenue for development of Proteolysis-Targeting Chimeras (PROTACs). Here, we present novel crystal structures of the HUWE1, TRIP12, and DTX1 WWE domains in complex with PAR building blocks and their analogs, thus enabling a comprehensive analysis of the PAR binding site structural diversity. Furthermore, we introduce a versatile toolbox of biophysical and biochemical assays for the discovery and characterization of novel WWE domain binders, including fluorescence polarization-based PAR binding and displacement assays, 15 N-NMR-based binding affinity assays and 19 F-NMR-based competition assays. Through these assays, we have characterized the binding of monomeric iso -ADP-ribose ( iso -ADPr) and its nucleotide analogs with the aforementioned WWE proteins. Finally, we have utilized the assay toolbox to screen a small molecule fragment library leading to the successful discovery of novel ligands targeting the HUWE1 WWE domain.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

Divergent viral phosphodiesterases for immune signaling evasion

Cyclic dinucleotides (CDNs) and other short oligonucleotides play fundamental roles in immune system activation in organisms ranging from bacteria to humans. In response, viruses use phosphodiesterase (PDE)-mediated oligonucleotide cleavage for immune evasion, a strategy whose diversity has not yet been explored. Here, we use a canonical 2H PDE (2H PDE) structure-based search of prokaryotic and eukaryotic viral sequences to identify an exceptional diversity of 2H PDEs across the virome, including enzymes not detectable with sequence search methods alone. Despite active site conservation, biochemical experiments reveal remarkable substrate specificity of these PDEs that corresponds to variations in the core 2H fold. This nuanced specificity allows 2H PDEs to selectively degrade oligonucleotide messengers to avoid interfering with host nucleotide signaling. Together, these findings nominate viral 2H PDEs as key regulators of CDN signaling across the tree of life.

CBASS↗

Full-Length ASFV B646L Gene Sequencing by Nanopore Offers a Simple and Rapid Approach for Identifying ASFV Genotypes

African swine fever (ASF) is an acute, highly hemorrhagic viral disease in domestic pigs and wild boars. The disease is caused by African swine fever virus, a double stranded DNA virus of the Asfarviridae family. ASF can be classified into 25 different genotypes, based on a 478 bp fragment corresponding to the C-terminal sequence of the B646L gene, which is highly conserved among strains and encodes the major capsid protein p72. The C-terminal end of p72 has been used as a PCR target for quick diagnosis of ASF, and its characterization remains the first approach for epidemiological tracking and identification of the origin of ASF in outbreak investigations. Recently, a new classification of ASF, based on the complete sequence of p72, reduced the 25 genotypes into only six genotypes; therefore, it is necessary to have the capability to sequence the full-length B646L gene (p72) in a rapid manner for quick genotype characterization. Here, we evaluate the use of an amplicon approach targeting the whole B646L gene, coupled with nanopore sequencing in a multiplex format using Flongle flow cells, as an easy, low cost, and rapid method for the characterization and genotyping of ASF in real-time.

Virology↗

A genomic perspective on fungal diversity and evolution

Originating from aquatic unicellular ancestors, over the course of ~1 billion years, the fungi have evolved to occupy nearly all aerobic environments on the planet, diversified into millions of different ‘species’ and have developed complex multicellular structures. Their relatively small, simple genomes have facilitated massive-scale sequencing and allowed us to explore genome evolution across an ancient eukaryotic kingdom. With thousands of genomes from diverse lineages now available, this Review will discuss insights into fungal biology and evolution gleaned with genomics and other multi-omics approaches. Using published genomes available through GenBank and the Joint Genome Institute’s MycoCosm platform, we generated kingdom-wide phylogenies and used them to highlight how fungal genomes have changed over time. With this phylogeny as a guide, we also discuss major evolutionary transitions that occurred across the fungal kingdom. Although progress has been made, these efforts are hampered by biases in genome representation and limited characterization of gene functions. Here, in this study, we discuss these challenges and possible future directions to address them, including initiatives to characterize conserved genes of unknown function and scale up sequencing towards 10,000 annotated fungal genomes.

Mondo, Stephen J. [USDOE Joint Genome Institute (J↗

Surfactant-like peptide gels are based on cross-β amyloid fibrils

Surfactant-like peptides, in which hydrophilic and hydrophobic residues are encoded within different domains in the peptide sequence, undergo facile self-assembly in aqueous solution to form supramolecular hydrogels. These peptides have been explored extensively as substrates for the creation of functional materials since a wide variety of amphipathic sequences can be prepared from commonly available amino acid precursors. The self-assembly behavior of surfactant-like peptides has been compared to that observed for small molecule amphiphiles in which nanoscale phase separation of the hydrophobic domains drives the self-assembly of supramolecular structures. Here, we investigate the relationship between sequence and supramolecular structure for a pair of bola-amphiphilic peptides, Ac-KLIIIK-NH 2 (L2) and Ac-KIIILK-NH 2 (L5). Despite similar length, composition, and polar sequence pattern, L2 and L5 form morphologically distinct assemblies, nanosheets and nanotubes, respectively. Cryo-EM helical reconstruction was employed to determine the structure of the L5 nanotube at near-atomic resolution. Rather than displaying self-assembly behavior analogous to conventional amphiphiles, the packing arrangement of peptides in the L5 nanotube displayed steric zipper interfaces that resembled those observed in the structures of β-amyloid fibrils. Like amyloids, the supramolecular structures of the L2 and L5 assemblies were sensitive to conservative amino acid substitutions within an otherwise identical amphipathic sequence pattern. This study highlights the need to better understand the relationship between sequence and supramolecular structure to facilitate the development of functional peptide-based materials for biomaterials applications.

Das, Abhinaba [Emory University, Atlanta, GA (Unit↗

Naturally ornate RNA-only complexes revealed by cryo-EM

The structures of natural RNAs remain poorly characterized and may hold numerous surprises. Here we report three-dimensional structures of three large ornate bacterial RNAs using cryo-electron microscopy (cryo-EM). GOLLD (Giant, Ornate, Lake- and Lactobacillales-Derived), ROOL (Rumen-Originating, Ornate, Large) and OLE (Ornate Large Extremophilic) RNAs form homo-oligomeric complexes whose stoichiometries are retained at lower concentrations than measured in cells. OLE RNA forms a dimeric complex with long co-axial pipes spanning two monomers. Both GOLLD and ROOL form distinct RNA-only multimeric nanocages with diameters larger than the ribosome, each empty except for a disordered loop. Extensive intramolecular and intermolecular A-minor interactions, kissing loops, an unusual A–A helix and other interactions stabilize the three complexes. Sequence covariation analysis of these large RNAs reveals evolutionary conservation of intermolecular interactions, supporting the biological importance of large, ornate RNA quaternary structures that can assemble without any involvement of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Cholesterol-dependent enzyme activity of human TSPO1

The amino acid sequence of the tryptophan-rich sensory proteins (TSPO) is substantially conserved throughout all kingdoms of life. Human mitochondrial TSPO1 (HsTSPO1) binds to porphyrins and steroids, although its interactions with these molecules remains unknown.HsTSPO1 is associated with numerous physiological and pathological disorders, but the underlying molecular mechanisms are unknown. Here, we disclose the finding of human mitochondrial TSPO as a cholesterol-dependent protoporphyrin IX oxygenase. The results of our biochemical characterization are consistent with structural data and evolutionary analysis. The dependence ofHsTSPO1 activity on cholesterol may be the result of the coevolution of this membrane protein with the membrane system. Our study provides a molecular foundation for comprehending the various roles played by mitochondrial TSPO in normal physiological and pathological situations.

Science & Technology - Other Topics↗

Borg extrachromosomal elements of methane-oxidizing archaea have conserved and expressed genetic repertoires

Borgs are huge extrachromosomal elements (ECE) of anaerobic methane-consuming “Candidatus Methanoperedens” archaea. Here, we used nanopore sequencing to validate published complete genomes curated from short reads and to reconstruct new genomes. 13 complete and four near-complete linear genomes share 40 genes that define a largely syntenous genome backbone. We use these conserved genes to identify new Borgs from peatland soil and to delineate Borg phylogeny, revealing two major clades. Remarkably, Borg genes encoding nanowire-like electron-transferring cytochromes and cell surface proteins are more highly expressed than those of host Methanoperedens, indicating that Borgs augment the Methanoperedens activity in situ. We reconstructed the first complete 4.00 Mbp genome for a Methanoperedens that is inferred to be a Borg host and predicted its methylation motifs, which differ from pervasive TC and CC methylation motifs of the Borgs. Thus, methylation may enable Methanoperedens to distinguish their genomes from those of Borgs. Very high Borg to Methanoperedens ratios and structural predictions suggest that Borgs may be capable of encapsulation. The findings clearly define Borgs as a distinct class of ECE with shared genomic signatures, establish their diversification from a common ancestor with genetic inheritance, and raise the possibility of periodic existence outside of host cells.

59 BASIC BIOLOGICAL SCIENCES↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

Comparative proteomics of a versatile, marine, iron-oxidizing chemolithoautotroph

This study conducted a comparative proteomic analysis to identify potential genetic markers for the biological function of chemolithoautotrophic iron oxidation in the marine bacterium Ghiorsea bivora. To date, this is the only characterized species in the class Zetaproteobacteria that is not an obligate iron-oxidizer, providing a unique opportunity to investigate differential protein expression to identify key genes involved in iron-oxidation at circumneutral pH. Over 1000 proteins were identified under both iron- and hydrogen-oxidizing conditions, with differentially expressed proteins found in both treatments. Notably, a gene cluster upregulated during iron oxidation was identified. This cluster contains genes encoding for cytochromes that share sequence similarity with the known iron-oxidase, Cyc2. Interestingly, these cytochromes, conserved in both Bacteria and Archaea, do not exhibit the typical β-barrel structure of Cyc2. This cluster potentially encodes a biological nanowire-like transmembrane complex containing multiple redox proteins spanning the inner membrane, periplasm, outer membrane, and extracellular space. The upregulation of key genes associated with this complex during iron-oxidizing conditions was confirmed by quantitative reverse transcription-PCR. These findings were further supported by electromicrobiological methods, which demonstrated negative current production by G. bivora in a three-electrode system poised at a cathodic potential. This research provides significant insights into the biological function of chemolithoautotrophic iron oxidation.

59 BASIC BIOLOGICAL SCIENCES↗

Conformational heterogeneity in the dGsw purine riboswitch: role of Mg²⁺ and 2’-dG in aptamer folding

Recent advancements in RNA structural biology have focused on unraveling the complexities of non-coding mRNA elements like riboswitches. These cis-acting regulatory regions undergo structural changes in response to specific cellular metabolites, leading to up or downregulation of downstream genes. The purine riboswitch family regulates many prokaryotic genes involved in purine degradation and biosynthesis. They feature an aptamer domain organized around a 3-way helical junction, where ligand encapsulation occurs at the junctional core. In our study, we chemically probed the aptamer domain of the 2’-dG-sensing purine riboswitch from Mesoplasma florum (dGsw) under various solution conditions to understand how Mg²⁺ and 2’-dG influence riboswitch folding. Here, we find that efficient 2’-dG binding strongly depends on Mg²⁺, indicating that Mg²⁺ is essential for priming dGsw for ligand interactions. We identified a previously undescribed sequence in the 5’ tail of dGsw that is complementary to a conserved helix. The inclusion of this region in a construct led to intramolecular competition between the alternate helix, Palt, and P1. Mutational analysis confirmed that 5’ flanking end of the aptamer domain forms an alternate helix in the absence of ligand. Molecular dynamics simulations revealed that this alternative conformation is stable. This helix may, therefore, facilitate the formation of an anti-terminator helix by opening the 3-way junction surrounding the 2’-dG binding site. Our study further establishes the importance of a closed terminal P1 helix conformation for metabolite binding and suggests that the delicate interplay between P1 and Palt may fine-tune downstream gene regulation. These insights offer a new perspective on riboswitch structure and enhance our understanding of the role that a conformational ensemble plays in riboswitch activity and regulation.

Biochemistry & Molecular Biology↗

A conserved viral RNA fold enables nuclease resistance across kingdoms of life

Abstract Viral exoribonuclease-resistant RNA (xrRNA) structures block cellular nucleases to produce subgenomic viral RNAs during infection. High sequence variability among xrRNAs from distantly related viruses raises questions about the shared molecular features that enable these RNAs to withstand the strong unwinding forces of exoribonucleases. Here, we present the first structure of a plant-virus xrRNA in its active conformation and uncover universal principles of xrRNA folding. Comparison with the structure of a human-pathogenic flavivirus xrRNA reveals that both share a core structural motif—a protective ring encircling the RNA’s 5′ end—despite lacking sequence similarity. Disrupting this core motif through targeted mutagenesis eliminates exoribonuclease-resistance and attenuates viral infection. We identify hundreds of related structures across multiple virus families, supporting the conservation of this mechanism. Our study demonstrates how distantly related RNA viruses have converged on a common structural strategy to inhibit cellular nucleases, with a universal ring topology as the defining feature of viral xrRNAs.

Biochemistry & Molecular Biology↗

Salt supplementation-induced metabolic reprogramming in Streptomyces coelicolor

Members of the genus Streptomyces are major producers of a wide variety of secondary metabolites that serve as bioactive compounds. Many secondary metabolites are produced in response to environmental signals such as biotic and abiotic stresses. In this study, we identified salt supplementation as one of the stimuli activating secondary metabolism in the model Streptomyces species, Streptomyces coelicolor. Comparative metabolomics revealed overproduction of several known secondary metabolites, most notably undecylprodigiosin and coelimycin P1, in addition to their biosynthetic intermediates and derivatives, as well as many unknown metabolites. Transcriptomic analysis revealed activation of diverse biological processes including cation uptake, compatible solute production, and the phosphate limitation stress response through conserved and species-specific mechanisms, presumably to overcome the increased salinity. This response leads to activation of a variety of regulatory and metabolic pathways required for production of secondary metabolites including activation of conserved metabolic pathways for energy and substrate supply and species-specific secondary metabolite biosynthetic gene clusters. Furthermore, several promoter sequences contributing to upregulation of secondary metabolism induced by salt supplementation were identified. Overall, our data show how S. coelicolor copes with the increased salinity and tailors the cellular metabolism toward secondary metabolism in a conserved and species-specific manner.

Otani, Hiroshi [USDOE Joint Genome Institute (JGI)↗

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES↗