Search NASA⌕ Search

SEARCH · Search NASA

Results for “Genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

Evolutionary history of arbuscular mycorrhizal fungi and genomic signatures of obligate symbiosis

The colonization of land and the diversification of terrestrial plants is intimately linked to the evolutionary history of their symbiotic fungal partners. Extant representatives of these fungal lineages include mutualistic plant symbionts, the arbuscular mycorrhizal (AM) fungi in Glomeromycota and fine root endophytes in Endogonales (Mucoromycota), as well as fungi with saprotrophic, pathogenic and endophytic lifestyles. These fungal groups separate into three monophyletic lineages but their evolutionary relationships remain enigmatic confounding ancestral reconstructions. Their taxonomic ranks are currently fluid. In this study, we recognize these three monophyletic linages as phyla, and use a balanced taxon sampling and broad taxonomic representation for phylogenomic analysis that rejects a hard polytomy and resolves Glomeromycota as sister to a clade composed of Mucoromycota and Mortierellomycota. Low copy numbers of genes associated with plant cell wall degradation could not be assigned to the transition to a plant symbiotic lifestyle but appears to be an ancestral phylogenetic signal. Both plant symbiotic lineages, Glomeromycota and Endogonales, lack numerous thiamine metabolism genes but the lack of fatty acid synthesis genes is specific to AM fungi. Many genes previously thought to be missing specifically in Glomeromycota are either missing in all analyzed phyla, or in some cases, are actually present in some of the analyzed AM fungal lineages, e.g. the high affinity phosphorus transporter Pho89. Based on a broad taxon sampling of fungal genomes we present a well-supported phylogeny for AM fungi and their sister lineages. We show that among these lineages, two independent evolutionary transitions to mutualistic plant symbiosis happened in a genomic background profoundly different from that known from the emergence of ectomycorrhizal fungi in Dikarya. These results call for further reevaluation of genomic signatures associated with plant symbiosis.

59 BASIC BIOLOGICAL SCIENCES↗

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Abstract Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcript’s impacts on plant metabolism has employed data from controlled environments. Here, we generated and analyzed data on the variation in the abundance of 26 metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies (GWAS)) and transcripts (using both transcriptome-wide association studies (TWAS) and an explainable artificial intelligence (AI) approach based on random forest (RF)) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

Mathivanan, Ramesh Kanna↗

Metagenome-assembled genomes provide insight into the metabolic potential during early production of Hydraulic Fracturing Test Site 2 in the Delaware Basin

Demand for natural gas continues to climb in the United States, having reached a record monthly high of 104.9 billion cubic feet per day (Bcf/d) in November 2023. Hydraulic fracturing, a technique used to extract natural gas and oil from deep underground reservoirs, involves injecting large volumes of fluid, proppant, and chemical additives into shale units. This is followed by a “shut-in” period, during which the fracture fluid remains pressurized in the well for several weeks. The microbial processes that occur within the reservoir during this shut-in period are not well understood; yet, these reactions may significantly impact the structural integrity and overall recovery of oil and gas from the well. To shed light on this critical phase, we conducted an analysis of both pre-shut-in material alongside production fluid collected throughout the initial production phase at the Hydraulic Fracturing Test Site 2 (HFTS 2) located in the prolific Wolfcamp formation within the Permian Delaware Basin of west Texas, USA. Specifically, we aimed to assess the microbial ecology and functional potential of the microbial community during this crucial time frame. Prior analysis of 16S rRNA sequencing data through the first 35 days of production revealed a strong selection for a Clostridia species corresponding to a significant decrease in microbial diversity. Here, we performed a metagenomic analysis of produced water sampled on Day 33 of production. This analysis yielded three high-quality metagenome-assembled genomes (MAGs), one of which was a Clostridia draft genome closely related to the recently classified Petromonas tenebris. This draft genome likely represents the dominant Clostridia species observed in our 16S rRNA profile. Annotation of the MAGs revealed the presence of genes involved in critical metabolic processes, including thiosulfate reduction, mixed acid fermentation, and biofilm formation. These findings suggest that this microbial community has the potential to contribute to well souring, biocorrosion, and biofouling within the reservoir. Our research provides unique insights into the early stages of production in one of the most prolific unconventional plays in the United States, with important implications for well management and energy recovery.

natural gas↗

Genomic Analysis of the Natural Variation of Fatty Acid Composition in Seed Oils of Camelina sativa

Camelina sativa is an oilseed crop that has shown strong promise as a biofuel feedstock. The profile of fatty acids greatly influences the oil quality; however, genetic mechanisms that determine the natural variation of fatty acid composition in camelina are not fully understood. A genome wide association study (GWAS) was performed to uncover genetic loci that may contribute to the contents of major fatty acids such as oleic and linolenic acids in camelina seed. Two approaches were taken to improve the GWAS efficiency. First, growing a diversity panel of 212 accessions in four locations and two nitrogen fertilization conditions revealed great variation in fatty acid contents in seeds. Second, using an improved reference genome, abundant markers, including 203,320 single nucleotide polymorphisms (SNPs) and 99,067 insertions/deletions (indels), were developed, which refined the population structure of the diversity panel. GWAS resulted in 118 genetic markers across 31 trait/treatment conditions. Closely linked markers were determined based on linkage decay and by comparing secondarily associated markers when highly associated ones were removed. Candidate genes were examined by comparing the pangenomes of 12 high-quality reference genomes. This study provides new resources to understand seed lipid metabolism and improve camelina oils through molecular breeding.

Life Sciences & Biomedicine - Other Topics↗

Status on Genetic Resistance to Rice Blast Disease in the Post-Genomic Era

Rice blast, caused by Magnaporthe oryzae, is a major threat to global rice production, necessitating the development of resistant cultivars through genetic improvement. Breakthroughs in rice genomics, including the complete genome sequencing of japonica and indica subspecies and the availability of various sequence-based molecular markers, have greatly advanced the genetic analysis of blast resistance. To date, approximately 122 blast-resistance genes have been identified, with 39 of these genes cloned and molecularly characterized. The application of these findings in marker-assisted selection (MAS) has significantly improved rice breeding, allowing for the efficient integration of multiple resistance genes into elite cultivars, enhancing both the durability and spectrum of resistance. Pangenomic studies, along with AI-driven tools like AlphaFold2, RoseTTAFold, and AlphaFold3, have further accelerated the identification and functional characterization of resistance genes, expediting the breeding process. Future rice blast disease management will depend on leveraging these advanced genomic and computational technologies. Emphasis should be placed on enhancing computational tools for the large-scale screening of resistance genes and utilizing gene editing technologies such as CRISPR-Cas9 for functional validation and targeted resistance enhancement and deployment. These approaches will be crucial for advancing rice blast resistance, ensuring food security, and promoting agricultural sustainability.

Pedrozo, Rodrigo↗

Ecological and genomic variation in ectomycorrhizal fungal exploration types

Ectomycorrhizal fungi (EMF) produce mycelia with variable extension and complexity, which can be classified according to soil ‘exploration types’ (ETs). ETs have received attention as one of the few mycorrhizal trait frameworks, but without an empirical classification of ET functional diversity and environmental preferences, understanding and interpreting EMF biogeographic patterns has been difficult. We conducted a synthesis combining: comparative EMF genomics to describe functional divergence in decomposition and nutrient cycling genes across ETs; and EMF trait distribution modeling across continental Europe, pairing soil and root EMF surveys to establish biogeographic ET niche profiles. We demonstrate a signature of ETs encoded in EMF genomes, which is independent from phylogeny and linked to biomass production strategies. EMF ET relative abundances were separated by soil, root, and dominant tree leaf type habitats and exhibited unique correlations with forest biotic (e.g. plant productivity and plant pathogen densities) and abiotic (e.g. nitrogen deposition and soil pH) conditions. These findings support a theory that EMF niche partitioning can be partially explained by extraradical mycelial traits, with underlying variation in ET biogeography likely arising from distinct decomposition and nutrient cycling potentials. We also identify important limitations to this trait framework and provide a guided outlook for future research.

biogeography↗

Characterisation and comparative analysis of mitochondrial genomes of false, yellow, black and blushing morels provide insights on their structure and evolution

Morchella species have considerable significance in terrestrial ecosystems, exhibiting a range of ecological lifestyles along the saprotrophism-to-symbiosis continuum. However, the mitochondrial genomes of these ascomycetous fungi have not been thoroughly studied, thereby impeding a comprehensive understanding of their genetic makeup and ecological role. In this study, we analysed the mitogenomes of 30 Morchellaceae species, including yellow, black, blushing and false morels. These mitogenomes are either circular or linear DNA molecules with lengths ranging from 217 to 565 kbp and GC content ranging from 38% to 48%. Fifteen core protein-coding genes, 28–37 tRNA genes and 3–8 rRNA genes were identified in these Morchellaceae mitogenomes. The gene order demonstrated a high level of conservation, with the cox1 gene consistently positioned adjacent to the rnS gene and cob gene flanked by apt genes. Some exceptions were observed, such as the rearrangement of atp6 and rps3 in Morchella importuna and the reversed order of atp6 and atp8 in certain morel mitogenomes. However, the arrangement of the tRNA genes remains conserved. We additionally investigated the distribution and phylogeny of homing endonuclease genes (HEGs) of the LAGLIDADG (LAGs) and GIY-YIG (GIYs) families. A total of 925 LAG and GIY sequences were detected, with individual species containing 19–48HEGs. These HEGs were primarily located in the cox1, cob, cox2 and nad5 introns and their presence and distribution displayed significant diversity amongst morel species. These elements significantly contribute to shaping their mitogenome diversity. Overall, this study provides novel insights into the phylogeny and evolution of the Morchellaceae.

59 BASIC BIOLOGICAL SCIENCES↗

Application of functional genomics for domestication of novel non-model microbes

Abstract With the expansion of domesticated microbes producing biomaterials and chemicals to support a growing circular bioeconomy, the variety of waste and sustainable substrates that can support microbial growth and production will also continue to expand. The diversity of these microbes also requires a range of compatible genetic tools to engineer improved robustness and economic viability. As we still do not fully understand the function of many genes in even highly studied model microbes, engineering improved microbial performance requires introducing genome-scale genetic modifications followed by screening or selecting mutants that enhance growth under prohibitive conditions encountered during production. These approaches include adaptive laboratory evolution, random or directed mutagenesis, transposon-mediated gene disruption, or CRISPR interference (CRISPRi). Although any of these approaches may be applicable for identifying engineering targets, here we focus on using CRISPRi to reduce the time required to engineer more robust microbes for industrial applications. One-Sentence Summary The development of genome scale CRISPR-based libraries in new microbes enables discovery of genetic factors linked to desired traits for engineering more robust microbial systems.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high-throughput identification and lineage assignment in Pakistan

Bacillus anthracis, the causative agent of anthrax, is a highly virulent zoonotic pathogen primarily affecting domesticated and wild herbivores. Human exposure to B. anthracis is primarily through contact with infected animals or contaminated animal products. In Pakistan, where livestock vaccines are largely unavailable and infected carcasses are often disposed of improperly, the risk to humans, wildlife and livestock is significant. Currently, the diagnosis of anthrax infections and outbreak tracing necessitates the isolation and culturing of B. anthracis, a process that requires BSL-3 facilities. In this study, we show that positive identification, genome reconstruction and lineage assignment can be accomplished using bioinformatic analysis of DNA extracted directly from environmental samples that would otherwise provide the starting material for isolation and culturing. This approach does not require laboratory target enrichment as is necessary for other pathogens, due in part to the extremely high bacterial load in the bloodstream in the deceased animals. Using these methods, we greatly expand the knowledge of endemic B. anthracis in Pakistan. We provide the first reference B. anthracis genomes from Pakistan since the 1970s and identify A.Br.014 Aust94 as a minor circulating sublineage alongside the dominant A.Br.047 Vollum. Future work will focus on the limits of detection and will determine if this bioinformatic method can be expanded more broadly for B. anthracis or other pathogens to replace typical culture-based methods.

A.Br.047 Vollum↗

Comparative physiological and genomic characterization of a novel Nitrobacter vulgaris strain from a nitrate-contaminated subsurface

Nitrite-oxidizing bacteria (NOB) represent a crucial node in the global nitrogen cycle. By catalyzing the second step of nitrification—the oxidation of nitrite to nitrate to generate energy for growth—NOB activity controls the fate of nitrite (NO 2 - ) in aerobic environments. Despite thriving in diverse environments, including soils, freshwater, marine ecosystems, subsurface habitats, and water treatment systems, organisms capable of nitrite oxidation are confined to Nitrobacter, Nitrospira, Nitrospina, Nitrotoga, and a few other specific lineages. The genus Nitrobacter, recognized for its facultative heterotrophic metabolism, is often associated with high-nitrogen environments. Here, we report the physiological characterization of a novel strain, Nitrobacter vulgaris strain MLSD-S22, isolated from a nitrate- and heavy-metal-contaminated subsurface. Growth inhibition experiments revealed that strain MLSD-S22 and the N. vulgaris type strain Z exhibited similar sensitivities to nitrite and nitrate, with nitrite being the most inhibitory. Microrespirometry demonstrated that the two N. vulgaris strains and Nitrobacter winogradskyi Nb-255 possessed higher affinities for nitrite and oxygen than previously reported for Nitrobacter, suggesting potential to compete in low-substrate environments. Long-read DNA sequencing provided a complete genome for strain MLSD-S22, revealing two plasmids and an intact nitrous oxide (N 2 O) reduction operon—an unexpected feature for Nitrobacter. While N 2 O reduction activity was not observed under the tested conditions, this discovery raises questions about the contribution of Nitrobacter NOB to the N 2 O sink. These findings broaden the physiological and genomic diversity of Nitrobacter, offering new insights into their adaptation strategies and providing a framework for future evaluation of their potential roles in nitrogen loss.

Nitrobacter↗

Phenotypic and genomic characterization of Methanothermobacter wolfeii strain BSEL, a CO 2 -capturing archaeon with minimal nutrient requirements

A new variant of Methanothermobacter wolfeii was isolated from an anaerobic digester using enrichment cultivation in anaerobic conditions. Here, the new isolate was taxonomically identified via 16S rRNA gene sequencing and tagged as M. wolfeii BSEL. The whole genome of the new variant was sequenced and de novo assembled. Genomic variations between the BSEL strain and the type strain were discovered, suggesting evolutionary adaptations of the BSEL strain that conferred advantages while growing under a low concentration of nutrients. M. wolfeii BSEL displayed the highest specific growth rate ever reported for the wolfeii species (0.27 ± 0.03 h –1 ) using carbon dioxide (CO 2 ) as unique carbon source and hydrogen (H 2 ) as electron donor. M. wolfeii BSEL grew at this rate in an environment with ammonium (NH 4 + ) as sole nitrogen source. The minerals content required to cultivate the BSEL strain was relatively low and resembled the ionic background of tap water without mineral supplements. Optimum growth rate for the new isolate was observed at 64°C and pH 8.3. In this work, it was shown that wastewater from a wastewater treatment facility can be used as a low-cost alternative medium to cultivate M. wolfeii BSEL. Continuous gas fermentation fed with a synthetic biogas mimic along with H 2 in a bubble column bioreactor using M. wolfeii BSEL as biocatalyst resulted in a CO 2 conversion efficiency of 97% and a final methane (CH 4 ) titer of 98.5%v, demonstrating the ability of the new strain for upgrading biogas to renewable natural gas.

09 BIOMASS FUELS↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Genome-scale Design and Engineering of Non-model Yeast Organisms for Production of Biofuels and Bioproducts

The overall goal of this project was to develop genome-scale design and engineering tools for two non-model yeast organisms including Rhodotorula toruloides and Issatchenkia orientalis to produce high-levels of fatty acids-derived products and organic acids, respectively. The project was performed between 9/15/2017 and 9/14/2024 (the last two-years were no-cost extensions). The team consisted of Huimin Zhao (Lead PI) and Christopher Rao (Co-PI) from the University of Illinois at Urbana-Champaign (UIUC), Costas Maranas (Co-PI) from the Pennsylvania State University, Joshua Rabinowitz (Co-PI) and Martin Wuhr (Co-PI) from Princeton University, and Yasuo Yoshikuni (Co-PI) from the DOE Joint Genome Institute. The team has made great progress in both tool development and fundamental understanding of these two non-model yeasts. In total, there were 40 research publications (one of them is still under review) and one patent application as well as numerous oral presentations.

60 APPLIED LIFE SCIENCES↗

Complete genome sequence of Sphingobium yanoikuyae strain CC4533

We have isolated a new strain of Sphingobium yanoikuyae , which belongs to the class Alphaproteobacteria, order Sphingomonadales, and family Sphingomonadaceae. This carotenoid-producing strain is capable of degrading xenobiotics and is tolerant to toxic levels of six heavy metals. We have designated the newly isolated strain of S. yanoikuyae as S. yanoikuyae strain CC4533 (hereafter called strain CC4533) because it was isolated from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii wild type strain CC4533. We sequenced the whole genome of strain CC4533 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of S. yanoikuyae strain CC4533 that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes of eight cultured microbes from soil sites in Wellesley, MA

We present the genomes of eight cultured microbes isolated from surface soil in Wellesley, MA. The dataset is useful for exploring genomic diversity among freshwater taxa including Pedobacter, Bacillus, Paenibacillus, Streptomyces, and Flavobacterium.

59 BASIC BIOLOGICAL SCIENCES↗

Genome_shuffling_enables_quantitative_trait_locus_mapping_in_Bacillus_subtilis

Genetic mapping is a powerful tool for eukaryotic genetics that has only been applied to bacteria in limited circumstances. Quantitative trait locus (QTL) mapping generally relies on sexual recombination to break linkages between genes, yet bacteria rarely undergo sufficient homologous recombination to generate suitable mapping populations. In this work, we used iterative biparental genome shuffling by protoplast fusion inBacillus subtilisto generate a population of bacteria with substantial random recombination throughout their genomes. Individual shuffled progeny were arrayed in well plates, resequenced, and characterized for a range of complex phenotypes including spore germination and swarming motility. Genetic mapping of the resulting phenotypes identified high-confidence QTLs of moderate size (∼10 kb), and these associations were validated through targeted genetic swaps. ThisB. subtilisQTL population can easily be used to map additional phenotypes, and the general approach for QTL mapping is applicable in a wide range of bacteria.

Bacillus subtilis↗