Search NASA⌕ Search

SEARCH · Search NASA

Results for “Genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Testing for the Genomic Footprint of Conflict Between Life Stages in an Angiosperm and Moss Species

Abstract The maintenance of genetic variation by balancing selection is of considerable interest to evolutionary biologists. An important but understudied potential driver of balancing selection is antagonistic pleiotropy between diploid and haploid stages of the plant life cycle. Despite sharing a common genome, sporophytes (2n) and gametophytes (n) may undergo differential or even opposing selection. Theoretical work suggests antagonistic pleiotropy between life stages can generate balancing selection and maintain genetic variation. Despite the potential for far-reaching consequences of gametophytic selection, empirical tests of its pleiotropic effects (neutral, synergistic, or antagonistic) on sporophytes are generally lacking. Here, we examined the population genomic signals of selection across life stages in the angiosperm Rumex hastatulus and the moss Ceratodon purpureus. We compared gene expression between life stages and sexes, combined with neutral diversity statistics and the analysis of the distribution of fitness effects. In contrast to what would be predicted under balancing selection due to antagonistic pleiotropy, we found that unbiased genes between life stages were under stronger purifying selection, likely explained by a predominance of synergistic pleiotropy between life stages and strong purifying selection on broadly expressed genes. In addition, we found that 30% of candidate genes under balancing selection in R. hastatulus were located within inversion polymorphisms. Our findings provide novel insights into the genome-wide characteristics and consequences of plant gametophytic selection.

Evolutionary Biology↗

Structural basis for a highly conserved RNA-mediated enteroviral genome replication

Abstract Enteroviruses contain conserved RNA structures at the extreme 5′ end of their genomes that recruit essential proteins 3CD and PCBP2 to promote genome replication. However, the high-resolution structures and mechanisms of these replication-linked RNAs (REPLRs) are limited. Here, we determined the crystal structures of the coxsackievirus B3 and rhinoviruses B14 and C15 REPLRs at 1.54, 2.2 and 2.54 Å resolution, revealing a highly conserved H-type four-way junction fold with co-axially stacked sA-sD and sB-sC helices that are stabilized by a long-range A•C•U base-triple. Such conserved features observed in the crystal structures also allowed us to predict the models of several other enteroviral REPLRs using homology modeling, which generated models almost identical to the experimentally determined structures. Moreover, our structure-guided binding studies with recombinantly purified full-length human PCBP2 showed that two previously proposed binding sites, the sB-loop and 3′ spacer, reside proximally and bind a single PCBP2. Additionally, the DNA oligos complementary to the 3′ spacer, the high-affinity PCBP2 binding site, abrogated its interactions with enteroviral REPLRs, suggesting the critical roles of this single-stranded region in recruiting PCBP2 for enteroviral genome replication and illuminating the promising prospects of developing therapeutics against enteroviral infections targeting this replication platform.

Biochemistry & Molecular Biology↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27↗

Integration of ultra-low coverage whole-genome sequences for reconstructing the evolutionary history of Galapagos giant tortoises

Genomic data from contemporary and historical samples often need to be coupled for evolutionary reconstructions of multitaxon complexes. However, the genetic data recovered from historical samples may result only in ultra-low coverage whole-genome sequences (ulcWGS; <0.15× depth), leading to inaccurate evolutionary inferences given a preponderance of missing data. Using the Galapagos giant tortoise radiation as a study system (Chelonoidis spp., composed of 13 extant and four extinct lineages), we assembled a novel methodological pipeline that removes potential noise introduced by the missing data and enhances the evolutionary signal from ulcWGS samples. We leveraged existing tools for phylogenomic placement (EPA-ng), population genomic structure (smartsnp) and admixture (Admixfrog, NGSadmix) to demonstrate that the evolutionary history of samples can be uncovered with sequencing depths as low as 0.008–0.139×. Importantly, these approaches do not use genotype imputation of the ulcWGS samples, which would require extensive reference datasets. Our application to two cases of extinct lineages of Galapagos giant tortoises, with and without references from the same lineage, demonstrates the general value of the approach. We confirm where the extinct lineages from San Cristóbal and Santa Fe islands fit into the Galapagos giant tortoise radiation, and that these lineages were evolutionarily distinct entities.

ancient DNA↗

Identification of proteins influencing CRISPR-associated transposases for enhanced genome editing

CRISPR-associated transposases (CASTs) hold tremendous potential for microbial genome editing because of their ability to integrate large DNA cargos in a programmable, site-specific manner. However, their widespread application has been hindered by poorly understood host factor requirements for transposition. To address this gap, we conducted the first genome-wide screen for host factors affecting Vibrio cholerae CAST (VchCAST) activity using an Escherichia coli RB-TnSeq library and identified 15 genes affecting VchCAST transposition. Of these, seven factors were validated to improve VchCAST activity, and two were inhibitory. Guided by the identification of homologous recombination effectors, RecD and RecA, we tested the λ-Red recombineering system in our VchCAST editing vectors and increased editing efficiency by 55.2-fold in E. coli, 5.6-fold in Pseudomonas putida, and 10.8-fold in Klebsiella michiganensis while maintaining high target specificity and similar insertion arrangements. This study improves the understanding of factors affecting VchCAST activity and enhances its efficiency as a bacterial genome editor.

Song, Leo C T↗

Genomic analysis and identification of a novel superantigen, SargEY, in Staphylococcus argenteus isolated from atopic dermatitis lesions

During surveillance of Staphylococcus aureus in lesions from patients with atopic dermatitis (AD), we isolated Staphylococcus argenteus, a species registered in 2011 as a new member of the genus Staphylococcus and previously considered a lineage of S. aureus. Genome sequence comparisons between S. argenteus isolates and representative S. aureus clinical isolates from various origins revealed that the S. argenteus genome from AD patients closely resembles that of S. aureus causing skin infections. We previously reported that 17%–22% of S. aureus isolated from skin infections produce staphylococcal enterotoxin Y (SEY), which predominantly induces T-cell proliferation via the T-cell receptor (TCR) Vα pathway. Complete genome sequencing of S. argenteus isolates revealed a gene encoding a protein similar to superantigen SEY, designated as SargEY, on its chromosome. Population structure analysis of S. argenteus revealed that these isolates are ST2250 lineage, which was the only lineage positive for the SEY-like gene among S. argenteus. Recombinant SargEY demonstrated immunological cross-reactivity with anti-SEY serum. SargEY could induce proliferation of human CD4 + and CD8 + T cells, as well as production of TNF-α and IFN-γ. SargEY showed emetic activity in a marmoset monkey model. S arg EY and SET (a phylogenetically close but uncharacterized SE) revealed their dependency on TCR Vα in inducing human T-cell proliferation. Additionally, TCR sequencing revealed other previously undescribed Vα repertoires induced by SEH. S arg EY and SEY may play roles in exacerbating the respective toxin-producing strains in AD.

59 BASIC BIOLOGICAL SCIENCES↗

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew↗

ATCCfinder - Download and Search the ATCC Genome Portal

Much strain-specific sequence data exists in research conducted before the deployment of large sequencing repositories, making it challenging to identify and validate the identity of strains used in these studies through bioinformatics and phenotyping. The American Type Culture Collection (ATCC) is an organization that sells a wide variety of microbes with strain-level taxonomy classification and associated sequenced reference genomes. Currently, ATCC does not provide a method for searching for sequence similarity between a query sequence and their database of reference genomes. Here I propose the software ATCCfinder, which utilizes ATCC application interface software (API) to generate query-able databases from ATCC Genome resources.

Koehler, Samuel↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

Evolutionary history of arbuscular mycorrhizal fungi and genomic signatures of obligate symbiosis

The colonization of land and the diversification of terrestrial plants is intimately linked to the evolutionary history of their symbiotic fungal partners. Extant representatives of these fungal lineages include mutualistic plant symbionts, the arbuscular mycorrhizal (AM) fungi in Glomeromycota and fine root endophytes in Endogonales (Mucoromycota), as well as fungi with saprotrophic, pathogenic and endophytic lifestyles. These fungal groups separate into three monophyletic lineages but their evolutionary relationships remain enigmatic confounding ancestral reconstructions. Their taxonomic ranks are currently fluid. In this study, we recognize these three monophyletic linages as phyla, and use a balanced taxon sampling and broad taxonomic representation for phylogenomic analysis that rejects a hard polytomy and resolves Glomeromycota as sister to a clade composed of Mucoromycota and Mortierellomycota. Low copy numbers of genes associated with plant cell wall degradation could not be assigned to the transition to a plant symbiotic lifestyle but appears to be an ancestral phylogenetic signal. Both plant symbiotic lineages, Glomeromycota and Endogonales, lack numerous thiamine metabolism genes but the lack of fatty acid synthesis genes is specific to AM fungi. Many genes previously thought to be missing specifically in Glomeromycota are either missing in all analyzed phyla, or in some cases, are actually present in some of the analyzed AM fungal lineages, e.g. the high affinity phosphorus transporter Pho89. Based on a broad taxon sampling of fungal genomes we present a well-supported phylogeny for AM fungi and their sister lineages. We show that among these lineages, two independent evolutionary transitions to mutualistic plant symbiosis happened in a genomic background profoundly different from that known from the emergence of ectomycorrhizal fungi in Dikarya. These results call for further reevaluation of genomic signatures associated with plant symbiosis.

59 BASIC BIOLOGICAL SCIENCES↗

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Abstract Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcript’s impacts on plant metabolism has employed data from controlled environments. Here, we generated and analyzed data on the variation in the abundance of 26 metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies (GWAS)) and transcripts (using both transcriptome-wide association studies (TWAS) and an explainable artificial intelligence (AI) approach based on random forest (RF)) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

Mathivanan, Ramesh Kanna↗

Metagenome-assembled genomes provide insight into the metabolic potential during early production of Hydraulic Fracturing Test Site 2 in the Delaware Basin

Demand for natural gas continues to climb in the United States, having reached a record monthly high of 104.9 billion cubic feet per day (Bcf/d) in November 2023. Hydraulic fracturing, a technique used to extract natural gas and oil from deep underground reservoirs, involves injecting large volumes of fluid, proppant, and chemical additives into shale units. This is followed by a “shut-in” period, during which the fracture fluid remains pressurized in the well for several weeks. The microbial processes that occur within the reservoir during this shut-in period are not well understood; yet, these reactions may significantly impact the structural integrity and overall recovery of oil and gas from the well. To shed light on this critical phase, we conducted an analysis of both pre-shut-in material alongside production fluid collected throughout the initial production phase at the Hydraulic Fracturing Test Site 2 (HFTS 2) located in the prolific Wolfcamp formation within the Permian Delaware Basin of west Texas, USA. Specifically, we aimed to assess the microbial ecology and functional potential of the microbial community during this crucial time frame. Prior analysis of 16S rRNA sequencing data through the first 35 days of production revealed a strong selection for a Clostridia species corresponding to a significant decrease in microbial diversity. Here, we performed a metagenomic analysis of produced water sampled on Day 33 of production. This analysis yielded three high-quality metagenome-assembled genomes (MAGs), one of which was a Clostridia draft genome closely related to the recently classified Petromonas tenebris. This draft genome likely represents the dominant Clostridia species observed in our 16S rRNA profile. Annotation of the MAGs revealed the presence of genes involved in critical metabolic processes, including thiosulfate reduction, mixed acid fermentation, and biofilm formation. These findings suggest that this microbial community has the potential to contribute to well souring, biocorrosion, and biofouling within the reservoir. Our research provides unique insights into the early stages of production in one of the most prolific unconventional plays in the United States, with important implications for well management and energy recovery.

natural gas↗

Genomic Analysis of the Natural Variation of Fatty Acid Composition in Seed Oils of Camelina sativa

Camelina sativa is an oilseed crop that has shown strong promise as a biofuel feedstock. The profile of fatty acids greatly influences the oil quality; however, genetic mechanisms that determine the natural variation of fatty acid composition in camelina are not fully understood. A genome wide association study (GWAS) was performed to uncover genetic loci that may contribute to the contents of major fatty acids such as oleic and linolenic acids in camelina seed. Two approaches were taken to improve the GWAS efficiency. First, growing a diversity panel of 212 accessions in four locations and two nitrogen fertilization conditions revealed great variation in fatty acid contents in seeds. Second, using an improved reference genome, abundant markers, including 203,320 single nucleotide polymorphisms (SNPs) and 99,067 insertions/deletions (indels), were developed, which refined the population structure of the diversity panel. GWAS resulted in 118 genetic markers across 31 trait/treatment conditions. Closely linked markers were determined based on linkage decay and by comparing secondarily associated markers when highly associated ones were removed. Candidate genes were examined by comparing the pangenomes of 12 high-quality reference genomes. This study provides new resources to understand seed lipid metabolism and improve camelina oils through molecular breeding.

Life Sciences & Biomedicine - Other Topics↗

Status on Genetic Resistance to Rice Blast Disease in the Post-Genomic Era

Rice blast, caused by Magnaporthe oryzae, is a major threat to global rice production, necessitating the development of resistant cultivars through genetic improvement. Breakthroughs in rice genomics, including the complete genome sequencing of japonica and indica subspecies and the availability of various sequence-based molecular markers, have greatly advanced the genetic analysis of blast resistance. To date, approximately 122 blast-resistance genes have been identified, with 39 of these genes cloned and molecularly characterized. The application of these findings in marker-assisted selection (MAS) has significantly improved rice breeding, allowing for the efficient integration of multiple resistance genes into elite cultivars, enhancing both the durability and spectrum of resistance. Pangenomic studies, along with AI-driven tools like AlphaFold2, RoseTTAFold, and AlphaFold3, have further accelerated the identification and functional characterization of resistance genes, expediting the breeding process. Future rice blast disease management will depend on leveraging these advanced genomic and computational technologies. Emphasis should be placed on enhancing computational tools for the large-scale screening of resistance genes and utilizing gene editing technologies such as CRISPR-Cas9 for functional validation and targeted resistance enhancement and deployment. These approaches will be crucial for advancing rice blast resistance, ensuring food security, and promoting agricultural sustainability.

Pedrozo, Rodrigo↗

Archaea: From Genomics to Physiology and the Origin of Life

This document represents a report on a meeting about Archaea. The meeting had an unusually diversified mix of topics all related to Archaea highlighting their differences and similarities with other kingdoms of life. Thus, a large number of scientists from others areas of biology participated in this conference. One-third of the speakers (11 of 33) represented laboratories whose main interests have not been archaea and who have not previously participated in similar symposia or workshops. Thus, this symposium provided a unique opportunity for archaeal researchers to interact in a wider forum. Because of the broad range of topics covered, the conference also introduced many of the participants to new areas of archaeal research. The discussions of genomics, molecular mechanisms of transcription, metabolic pathways and evolution were at a very high level. Talks and posters provided detailed discussions of the state of the current knowledge in RNA processing, transcriptional initiation, chromatin structure, aminoacyl-tRNA synthetases, autotrophic CO2 fixation, Upid biosynthesis and a wide range of other topics. In addition to providing overviews, major areas of scientific argument were clearly delineated, particularly in the discussions of genomics and evolution. Some of the questions raised included: how representative are individual gene trees of organismal evolution, how prevalent is horizontal evolution, how reliable are functional assignments in genomics? On these topics, the different points of view were well represented. The future of any field depends on the enthusiasm and intellectual engagement of young scientists working in the area. Therefore, the participation of 29 graduate and postdoctoral students (out of about 135 participants) was a highlight of the meeting. This was the consequence of funding contributions by NSF and NASA.

Vothknecht, Ute C.↗