Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Transcription factor binding divergence drives transcriptional and phenotypic variation in maize

Regulatory elements are essential components of plant genomes that have shaped the domestication and improvement of modern crops. However, their identity, function and diversity remain poorly characterized, limiting our ability to harness their full power for agricultural advances using induced or natural variation. Here, in this study, we mapped transcription factor (TF) binding for 200 TFs from 30 families in two distinct maize inbred lines historically used in maize breeding. TF binding comparison revealed widespread differences between inbreds, driven largely by structural variation, that correlated with gene expression changes and explained complex quantitative trait loci such as Vgt1, an important determinant of flowering time, and DICE, an herbivore resistance enhancer. CRISPR–Cas9 editing of TF binding regions validated the function and structure of regulatory regions at various loci controlling plant architecture and biotic resistance. Our maize TF binding catalogue identifies functional regulatory regions and enables collective and comparative analysis, highlighting its value for agricultural improvement.

Galli, Mary [Rutgers Univ., Piscataway, NJ (United↗

Soil moisture and temperature from 2019 to 2024 along northeast- and southwest-facing hillslopes at the Lower Montane site in the East River Watershed, Colorado

Soil moisture, temperature, and electrical conductivity have been monitored at multiple depths (between 10 and 50 cm) at 4 locations along a northeast-facing slope and 3 locations on the opposite southwest-facing slope at the Lower Montane site in the East River Watershed, Colorado, from Oct 2019 to Oct 2024. The purpose of this data is to inform hydro-biogeochemical analyses for the Watershed Function Scientific Focus Area (SFA). Two locations on the northeast-facing slope were reinstalled in 2020 due to damage from wildlife, and thus data for these sites are provided in two distinct files. Overall, the data are reported in 9 CSV files containing the measurements, and the locations are provided in the Sensor_Location.csv file. There is a total of 10 *.csv data files and 3 *.csv metadata files. Older datasets associated with the northeast-facing slope are provided in another archive (see reference). These data products are part of the Watershed Function Scientific Focus Area collection effort to further scientific understanding of biogeochemical dynamics from genome to watershed scales. Feel free to contact the author with any questions or collaboration interests.

54 ENVIRONMENTAL SCIENCES↗

Pseudomonas aeruginosa gene PA4880 encodes a Dps-like protein with a Dps fold, bacterioferritin-type ferroxidase centers, and endonuclease activity

We report the biochemical, structural, and functional characterization of the protein coded by gene PA4880 in the P. aeruginosa PAO1 genome. The PA4880 gene had been annotated as coding a probable bacterioferritin. Our structural work shows that the product of gene PA4880 is a protein that adopts the Dps subunit fold, which oligomerizes into a 12-mer quaternary structure. Unlike Dps, however, the ferroxidase di-iron centers and iron coordinating ligands are buried within each subunit, in a manner identical to that observed in the ferroxidase center of P. aeruginosa bacterioferritin. Since these structural characteristics correspond to Dps-like proteins, we term the protein as P. aeruginosa Dps-like, or Pa DpsL. The ferroxidase centers in Pa DpsL catalyze the oxidation of Fe 2+ utilizing O 2 or H 2 O 2 as oxidant, and the resultant Fe 3+ is compartmentalized in the interior cavity. Interestingly, incubating Pa DpsL with plasmid DNA results in efficient nicking of the DNA and at higher concentrations of Pa DpsL the DNA is linearized and eventually degraded. The nickase and endonuclease activities suggest that Pa DpsL, in addition to participating in the defense of P. aeruginosa cells against iron-induced toxicity, may also participate in the innate immune mechanisms consisting of restriction endonucleases and cognate methyl transferases.

59 BASIC BIOLOGICAL SCIENCES↗

Saccule contribution to immediate early gene induction in the gerbil brainstem with posterior canal galvanic or hypergravity stimulation

Immunolabeling patterns of the immediate early gene-related protein Fos in the gerbil brainstem were studied following stimulation of the sacculus by both hypergravity and galvanic stimulation. Head-restrained, alert animals were exposed to a prolonged (1 h) inertial vector of 2 G (19.6 m/s2) head acceleration directed in a dorso-ventral head axis to maximally stimulate the sacculus. Fos-defined immunoreactivity was quantified, and the results compared to a control group. The hypergravity stimulus produced Fos immunolabeling in the dorsomedial cell column (dmcc) of the inferior olive independently of other subnuclei. Similar dmcc labeling was induced by a 30 min galvanic stimulus of up to -100 microA applied through a stimulating electrode placed unilaterally on the bony labyrinth overlying the posterior canal (PC). The pattern of vestibular afferent firing activity induced by this galvanic stimulus was quantified in anesthetized gerbils by simultaneously recording from Scarpa's ganglion. Only saccular and PC afferent neurons exhibited increases in average firing rates of 200-300%, suggesting a pattern of current spread involving only PC and saccular afferent neurons at this level of stimulation. These results suggest that alteration in saccular afferent firing rates are sufficient to induce Fos-defined genomic activation of the dmcc, and lend further evidence to the existence of a functional vestibulo-olivary-cerebellar pathway of adaptation to novel gravito-inertial environments.

NASA Discipline Neuroscience↗

Functional insights of novel Bathyarchaeia reveal metabolic versatility in their role in peatlands of the Peruvian Amazon

ABSTRACT The decomposition of soil organic carbon within tropical peatlands is influenced by the functional composition of the microbial community. In this study, building upon our previous work, we recovered a total of 28 metagenome-assembled genomes (MAGs) classified as Bathyarchaeia from the tropical peatlands of the Pastaza-Marañón Foreland Basin (PMFB) in the Amazon. Using phylogenomic analyses, we identified nine genus-level clades to have representatives from the PMFB, with four forming a putative novel family (“CandidatusPaludivitaceae”) endemic to peatlands. We focus on theCa. Paludivitaceae MAGs due to the novelty of this group and the limited understanding of their role within tropical peatlands. Functional analysis of these MAGs reveals that this putative family comprises facultative anaerobes, possessing the genetic potential for oxygen, sulfide, or nitrogen oxidation. This metabolic versatility can be coupled to the fermentation of acetoin, propanol, or proline. The other clades outsideCa. Paludivitaceae are putatively capable of acetogenesis andde novoamino acid biosynthesis and encode a high amount of Fe 3+ transporters. Crucially, theCa. Paludivitaceae are predicted to be carboxydotrophic, capable of utilizing CO for energy generation or biomass production. Through this metabolism, they could detoxify the environment from CO, a byproduct of methanogenesis, or produce methanogenic substrates like CO 2 and H 2 . Overall, our results show the complex metabolism and various lineages of Bathyarchaeia within tropical peatlands pointing to the need to further evaluate their role in these ecosystems. IMPORTANCE With the expansion of theCandidatusPaludivitaceae family by the assembly of 28 new metagenome assembled genomes, this study provides novel insights into their metabolic diversity and ecological significance in peatland ecosystems. From a comprehensive phylogenic and functional analysis, we have elucidated their putative unique facultative anaerobic capabilities and CO detoxification potential. This research highlights their crucial role in carbon cycling and greenhouse gas regulation. These findings are essential for resolving the microbial processes affecting peat soil stability, offering new perspectives on the ecological roles of previously underexplored and underrepresented archaeal populations.

Microbiology↗

An Inherited Efficiencies Model of Non-Genomic Evolution

A model for the evolution of biological systems in the absence of a nucleic acid-like genome is proposed and applied to model the earliest living organisms -- protocells composed of membrane encapsulated peptides. Assuming that the peptides can make and break bonds between amino acids, and bonds in non-functional peptides are more likely to be destroyed than in functional peptides, it is demonstrated that the catalytic capabilities of the system as a whole can increase. This increase is defined to be non-genomic evolution. The relationship between the proposed mechanism for evolution and recent experiments on self-replicating peptides is discussed.

New, Michael H.↗

Functional Horizontal Gene Transfer From Bacteria to Eukaryotes

The antiquity of bacteria and archaea in part explains why they, along with viruses, encode most of the genetic and biochemical diversity on Earth. Eukaryotic life evolved into a world teeming with prokaryotes, and so bacteria (especially) have inevitably affected eukaryotic biology as parasitic, commensal, or beneficial symbionts. But along with these important organismal interactions, the ubiquity and diversity of bacteria have also made them frequent sources of horizontally transferred DNA into eukaryotic genomes. Here we survey the role of bacterial genes throughout the eukaryotic lineage. We review what steps horizontal gene transfers (HGTs) take in becoming functional, what bacterial groups these HGTs come from, and what functions these HGTs typically bestow on their eukaryotic recipient. We classify HGTs into two broad types: those that maintain preexisting functions and those that add new functionality to the recipient. We find that genes involved in host nutrition, protection, and adaptation to extreme environments are the most common HGTs from bacterial to eukaryotic genomes.

Bacteria↗

GENTANGLE: integrated computational design of gene entanglements

The design of two overlapping genes in a microbial genome is an emerging technique for adding more reliable control mechanisms in engineered organisms for increased stability. The design of functional overlapping gene pairs is a challenging procedure, and computational design tools are used to improve the efficiency to deploy successful designs in genetically engineered systems. GENTANGLE (Gene Tuples ArraNGed in overLapping Elements) is a high-performance containerized pipeline for the computational design of two overlapping genes translated in different reading frames of the genome. This new software package can be used to design and test gene entanglements for microbial engineering projects using arbitrary sets of user-specified gene pairs.

59 BASIC BIOLOGICAL SCIENCES↗

Status on Genetic Resistance to Rice Blast Disease in the Post-Genomic Era

Rice blast, caused by Magnaporthe oryzae, is a major threat to global rice production, necessitating the development of resistant cultivars through genetic improvement. Breakthroughs in rice genomics, including the complete genome sequencing of japonica and indica subspecies and the availability of various sequence-based molecular markers, have greatly advanced the genetic analysis of blast resistance. To date, approximately 122 blast-resistance genes have been identified, with 39 of these genes cloned and molecularly characterized. The application of these findings in marker-assisted selection (MAS) has significantly improved rice breeding, allowing for the efficient integration of multiple resistance genes into elite cultivars, enhancing both the durability and spectrum of resistance. Pangenomic studies, along with AI-driven tools like AlphaFold2, RoseTTAFold, and AlphaFold3, have further accelerated the identification and functional characterization of resistance genes, expediting the breeding process. Future rice blast disease management will depend on leveraging these advanced genomic and computational technologies. Emphasis should be placed on enhancing computational tools for the large-scale screening of resistance genes and utilizing gene editing technologies such as CRISPR-Cas9 for functional validation and targeted resistance enhancement and deployment. These approaches will be crucial for advancing rice blast resistance, ensuring food security, and promoting agricultural sustainability.

Pedrozo, Rodrigo↗

A Trichomonas vaginalis C2-XYPPX-repeat protein with a structured C2 domain displaying dampened flexibility upon binding calcium

C2 domains are ubiquitous membrane-binding modules of ∼130 residues in eukaryotes that are often associated with proteins involved in membrane trafficking and lipid modification. The genome of Trichomonas vaginalis, the most common, non-viral, sexually transmitted human pathogen, encodes eight genes that contain a N-terminal C2 module linked to a XYPPX-repeat domain of more than four XYPPX repeats (C2-XYPPX). While the function of the XYPPX-repeat domain remains unknown, its multiple association with C2 domains in T. vaginalis suggests it is important. Here, the C2 domain from one of these C2-XYPPX-repeat proteins, Tv-C2-1, was structurally and physically characterized using X-ray crystallography and NMR spectroscopy. The crystal structure for Tv-C2-1 shows that this domain shares a fold common to all C2 domains, a compact Greek-key motif composed of eight anti-parallel β-strands in the type-2 topology. An NMR chemical shift perturbation study with Ca 2+ showed that Tv-C2-1 bound two Ca 2+ atoms primarily via two loops (loop-1 and loop-3) on the predicted calcium binding face of the protein with K d s of 58.0 ± 0.1 μM and 232 ± 6 μM. Estimations of the overall rotational correlation time, τ c , in the apo (11.1 ns) and Ca 2+ -bound (9.2 ns) state suggests the protein becomes more compact upon Ca 2+ binding, consistent with a decrease in dynamics in loop-3 and marginally in loop-1 suggested by amide 15 N heteronuclear steady-state { 1 H}- 15 N NOEs. Showing Tv-C2-1 binds calcium and adopts a compact Greek-key motif structure, two primary features of C2 domains, suggests understanding the function of the XYPPX-repeat domain may be warranted.

NMR spectroscopy↗

Atmospheric methane consumption in arid ecosystems acts as a reverse chimney and is accelerated by plant-methanotroph biomes

Drylands cover one-third of the Earth’s surface and are one of the largest terrestrial sinks for methane. Understanding the structure–function interplay between members of arid biomes can provide critical insights into mechanisms of resilience toward anthropogenic and climate-change-driven environmental stressors—water scarcity, heatwaves, and increased atmospheric greenhouse gases. This study integrates in situ measurements with culture-independent and enrichment-based investigations of methane-consuming microbiomes inhabiting soil in the Anza-Borrego Desert, a model arid ecosystem in Southern California, United States. The atmospheric methane consumption ranged between 2.26 and 12.73 μmol m 2 h −1 , peaking during the daytime at vegetated sites. Metagenomic studies revealed similar soil-microbiome compositions at vegetated and unvegetated sites, with Methylocaldum being the major methanotrophic clade. Eighty-four metagenome-assembled genomes were recovered, six represented by methanotrophic bacteria (three Methylocaldum , two Methylobacter , and uncultivated Methylococcaceae ). The prevalence of copper-containing methane monooxygenases in metagenomic datasets suggests a diverse potential for methane oxidation in canonical methanotrophs and uncultivated Gammaproteobacteria. Five pure cultures of methanotrophic bacteria were obtained, including four Methylocaldum . Genomic analysis of Methylocaldum isolates and metagenome-assembled genomes revealed the presence of multiple stand-alone methane monooxygenase subunit C paralogs, which may have functions beyond methane oxidation. Furthermore, these methanotrophs have genetic signatures typically linked to symbiotic interactions with plants, including tryptophan synthesis and indole-3-acetic acid production. Based on in situ fluxes and soil microbiome compositions, we propose the existence of arid-soil reverse chimneys, an empowered methane sink represented by yet-to-be-defined cooperation between desert vegetation and methane-consuming microbiomes.

59 BASIC BIOLOGICAL SCIENCES↗

An archaeal genomic signature

Comparisons of complete genome sequences allow the most objective and comprehensive descriptions possible of a lineage's evolution. This communication uses the completed genomes from four major euryarchaeal taxa to define a genomic signature for the Euryarchaeota and, by extension, the Archaea as a whole. The signature is defined in terms of the set of protein-encoding genes found in at least two diverse members of the euryarchaeal taxa that function uniquely within the Archaea; most signature proteins have no recognizable bacterial or eukaryal homologs. By this definition, 351 clusters of signature proteins have been identified. Functions of most proteins in this signature set are currently unknown. At least 70% of the clusters that contain proteins from all the euryarchaeal genomes also have crenarchaeal homologs. This conservative set, which appears refractory to horizontal gene transfer to the Bacteria or the Eukarya, would seem to reflect the significant innovations that were unique and fundamental to the archaeal "design fabric." Genomic protein signature analysis methods may be extended to characterize the evolution of any phylogenetically defined lineage. The complete set of protein clusters for the archaeal genomic signature is presented as supplementary material (see the PNAS web site, www.pnas.org).

Non-NASA Center↗

Phosphorylation toggles the SARS-CoV-2 nucleocapsid protein between two membrane-associated condensate states

Abstract The Nucleocapsid protein (N) of SARS-CoV-2 plays a critical role in the viral lifecycle by regulating RNA replication and by packaging the viral genome. N and RNA phase separate to form condensates that may be important for these functions. Both functions occur at membrane surfaces, but how N toggles between these two membrane-associated functional states is unclear. Here, we reveal that phosphorylation switches how N condensates interact with membranes, in part by modulating condensate material properties. Our studies also show that phosphorylation alters N’s interaction with viral membrane proteins. We gain mechanistic insight through structural analysis and molecular simulations, which suggest phosphorylation induces a conformational change in N that softens condensate material properties. Together, our findings identify membrane association as a key feature of N condensates and provide mechanistic insights into the regulatory role of phosphorylation. Understanding this mechanism suggests potential therapeutic targets for COVID infection.

Science & Technology - Other Topics↗

Transporter annotations are holding up progress in metabolic modeling

Mechanistic, constraint-based models of microbial isolates or communities are a staple in the metabolic analysis toolbox, but predictions about microbe-microbe and microbe-environment interactions are only as good as the accuracy of transporter annotations. A number of hurdles stand in the way of comprehensive functional assignments for membrane transporters. These include general or non-specific substrate assignments, ambiguity in the localization, directionality and reversibility of a transporter, and the many-to-many mapping of substrates, transporters and genes. In this perspective, we summarize progress in both experimental and computational approaches used to determine the function of transporters and consider paths forward that integrate both. Investment in accurate, high-throughput functional characterization is needed to train the next-generation of predictive tools toward genome-scale metabolic network reconstructions that better predict phenotypes and interactions. More reliable predictions in this domain will benefit fields ranging from personalized medicine to metabolic engineering to microbial ecology.

Casey, John↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

The evolution of transcriptional regulation in eukaryotes

Gene expression is central to the genotype-phenotype relationship in all organisms, and it is an important component of the genetic basis for evolutionary change in diverse aspects of phenotype. However, the evolution of transcriptional regulation remains understudied and poorly understood. Here we review the evolutionary dynamics of promoter, or cis-regulatory, sequences and the evolutionary mechanisms that shape them. Existing evidence indicates that populations harbor extensive genetic variation in promoter sequences, that a substantial fraction of this variation has consequences for both biochemical and organismal phenotype, and that some of this functional variation is sorted by selection. As with protein-coding sequences, rates and patterns of promoter sequence evolution differ considerably among loci and among clades for reasons that are not well understood. Studying the evolution of transcriptional regulation poses empirical and conceptual challenges beyond those typically encountered in analyses of coding sequence evolution: promoter organization is much less regular than that of coding sequences, and sequences required for the transcription of each locus reside at multiple other loci in the genome. Because of the strong context-dependence of transcriptional regulation, sequence inspection alone provides limited information about promoter function. Understanding the functional consequences of sequence differences among promoters generally requires biochemical and in vivo functional assays. Despite these challenges, important insights have already been gained into the evolution of transcriptional regulation, and the pace of discovery is accelerating.

Review, Academic↗