Search NASA⌕ Search

SEARCH · Search NASA

Results for “Functional Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Status of genome function annotation in model organisms and crops

Abstract Since the entry into genome‐enabled biology several decades ago, much progress has been made in determining, describing, and disseminating the functions of genes and their products. Yet, this information is still difficult to access for many scientists and for most genomes. To provide easy access and a graphical summary of the status of genome function annotation for model organisms and bioenergy and food crop species, we created a web application ( https://genomeannotation.rheelab.org ) to visualize, search, and download genome annotation data for 28 species. The summary graphics and data tables will be updated semi‐annually, and snapshots will be archived to provide a historical record of the progress of genome function annotation efforts. Clear and simple visualization of up‐to‐date genome function annotation status, including the extent of what is unknown, will help address the grand challenge of elucidating the functions of all genes in organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Structural models and functional annotations for the Sphagnum divinum proteome

This dataset contains the structural models for the primary transcripts of the Sphagnum divinum proteome. Additionally, for a subset of these proteins, sequence and structural alignment results are provided. This dataset represents the most thorough structural study of a Sphagnum species, also known as peat mosses, by providing three-dimensional atomic resolution structures of the majority of the encoded proteins as well as structural alignment results used in the application of annotating the proteome. References (DOI) AlphaFold v2 Monomer: https://doi.org/10.1038/s41586-021-03819-2. References (DOI) US-align2: https://doi.org/10.1038/s41592-022-01585-1

59 BASIC BIOLOGICAL SCIENCES↗

In–context promoter bashing of the Sorghum bicolor gene models functionally annotated as bundle sheath cell preferred expressing phosphoenolpyruvate carboxykinase and alanine aminotransferase

In-context promoter bashing via genome editing is a route to identify and characterize critical regulatory regions that govern expression of genes of interest. The outcomes of in-context promoter bashing can be used to inform editing strategies to modulate the expression of selected gene models in a desired fashion. Here, we employed in-context promoter bashing to characterize the proximal upstream regulatory regions of sorghum genes encoding phosphoenolpyruvate carboxykinase bundle sheath (SbPEPCK.BS, SbiTx430.01G455400) and alanine aminotransferase bundle sheath (SbAlaAT.BS, SbiTx430.02G006600), two proteins involved in the PCK C 4 pathway. Characterized germinal edits within the targeted regions upstream of these two genes ranged in size from 138 up to 1790 bp. A 138 bp within the SbPEPCK.BS upstream region and a 1643 bp element within the SbAlaAT.BS upstream region were determined to be important for maintenance of transcription levels. No change in development or various physiological parameters was observed in characterized lineages carrying promoter edits. However, significant changes in seed reserves and a reduction in 100-seed weight were consistently observed, under both greenhouse and field environments, in plants carrying an edit in the promoter of SbPEPCK.BS gene, which were significantly reduced in transcript accumulation for this gene.

60 APPLIED LIFE SCIENCES↗

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab↗

High-Throughput Directed Evolution of Marine Microalgae and Phototrophic Consortia for Improved Biomass Yields (Final Report)

Primary project achievements include using selective pressures (O 2 , light, temperature) and developing culturing regimes for the diatom Nitzschia inconspicua str. hildebrandi to attain enrichments with an ~90% increase in areal biomass productivity relative to the parental strain under pond-mimicking conditions with high O 2 stress in laboratory bioreactors. The resulting strain (GAI-337) was tested further for dilution time, culture density, CO 2 supplementation, pH, temperature, and dissolved O 2 concentration under outdoor pond-mimicking conditions to improve areal productivities. These experiments yielded an optimum harvest and dilution time just after sunset, ~0.45 g AFDW L -1 initial culture density for maximal productivities, no requirement for CO 2 supplementation or pH control, maximal performance under a diel temperature curve going from 24 °C at night to 36 °C during the day, and benefits from some O 2 removal from the culture by bubbling with air. Using pond-mimicking laboratory bioreactors, N. inconspicua GAI-337 achieved ~42 g AFDW m -2 d -1 . Nutrient limitation experiments resulted in a biomass composition that equated to ~160 Gallons of Gasoline Equivalent energy per ton AFDW, highlighting the potential of GAI-337 as a promising renewable fuel feedstock strain. Genome resequencing has revealed genome alterations potentially contributing to the improved growth of GAI-337 in the laboratory. Based on the comparative analyses of the GAI-337 and GAI-229 (reference) strains, we identified 144 single nucleotide substitutions that resulted in amino acid change, 7 single nucleotide substitutions that resulted in protein truncation; 5 deletions; and 1 frameshift mutation. From the mutations that potentially affect expression of functionally annotated genes, particular interest was noted for an interferon-induced 6-16 family protein that may be involved in the host immune response against microbe invasion; the chaperone protein DnaK, which may function to protect the folding of proteins within the cell; and SPRY domain protein that is found in many eukaryotic proteins important in cell signaling pathways. Transcriptome analysis revealed over 1000 genes with increased transcript levels. Many of these and many of the genes with mutations are not yet functionally annotated and an increased bioinformatics effort is necessary to more completely analyze the Nitzschia inconspicua genome. Adaptive laboratory evolution (ALE) was performed for over 300 days using consecutive 0.5°C temperature increases in a constant temperature incubator to attain greater thermal tolerance in Nitzschia inconspicua. The adapted strain was able to grow at a constant temperature of 37.5°C; whereas this constant temperature was lethal to the parental control, which had an upper temperature boundary of 35.5°C prior to adaptive evolution. Several high-temperature clonal isolates were obtained from the evolved population following ALE, and increased temperature tolerance was observed in clonal adapted cultures. The final temperature adaptation was maintained through cryopreservation and was observed in multiple clonal isolates, including multiple clonal isolates with significantly increased cell size, indicating the potential occurrence of a sexual cycle during the clonal isolation process. A survey of Nannochloropsis strains was conducted for tolerances to high pH and high bicarbonate media. Nannochloropsis granulata showed promising growth in diel bioreactors and was successfully grown at the GAI Kauai farm site in long-term growth campaigns. Co-culturing using Nitzschia inconspicua, Nannochloropsis and a cyanobacterium were assembled in the laboratory to determine if productivity synergies could be attained. Although all strains grew well in the laboratory high-bicarbonate media individually, the cyanobacterium quickly outgrew the other strains in the laboratory consortium pushing the co-culture away from a diverse (and potentially synergistic assemblage) phototroph culture towards a monoculture dominated by the cyanobacterium. Several outdoor growth campaigns were conducted, with productivities ranging between 10-20 g/m 2 /d of biomass. The best performing strain in the laboratory (GAI-337) did not outperform reference strains at the Kauai farm under the conditions used. Addition growth campaigns are necessary under conditions that result in higher biomass (>20 g/m 2 /d) and that attain higher O 2 levels are likely necessary. Initial data indicate that the thermally adapted strain did slightly better than the control strain at higher temperatures; however, additional campaigns are necessary to establish statistical significance. In summary, Nitzschia inconspicua is able to attain exemplary biomass and lipid yields in the laboratory bioreactors. Strain evolution to both O 2 and temperature resulted in targeted strain improvements. Additional outdoor campaigns are necessary to determine if laboratory improvements translate to the field.

09 BIOMASS FUELS↗

Tropical lacustrine sediment microbial community response to an extreme El Niño event

Salinity can influence microbial communities and related functional groups in lacustrine sediments, but few studies have examined temporal variability in salinity and associated changes in lacustrine microbial communities and functional groups. To better understand how microbial communities and functional groups respond to salinity, we examined geochemistry and functional gene amplicon sequence data collected from 13 lakes located in Kiritimati, Republic of Kiribati (2° N, 157° W) in July 2014 and June 2019, dates which bracket the very large El Niño event of 2015–2016 and a period of extremely high precipitation rates. Lake water salinity values in 2019 were significantly reduced and covaried with ecological distances between microbial samples. Specifically, phylum- and family-level results indicate that more halophilic microorganisms occurred in 2014 samples, whereas more mesohaline, marine, or halotolerant microorganisms were detected in 2019 samples. Functional Annotation of Prokaryotic Taxa (FAPROTAX) and functional gene results (nifH, nrfA, aprA) suggest that salinity influences the relative abundance of key functional groups (chemoheterotrophs, phototrophs, nitrogen fixers, denitrifiers, sulfate reducers), as well as the microbial diversity within functional groups. Accordingly, we conclude that microbial community and functional gene groups in the lacustrine sediments of Kiritimati show dynamic changes and adaptations to the fluctuations in salinity driven by the El Niño-Southern Oscillation.

59 BASIC BIOLOGICAL SCIENCES↗

DiMER

SAND2025-04145O DiMER is a Python based tool that helps researchers understand the functions of genes by searching through multiple biological databases. It takes user-provided data and scans various databases to find the best matches for gene functions, generating a clear summary of results. DiMER identifies the most relevant functional annotations and improves upon previous annotations by replacing instances of "unknown protein function" with more accurate descriptions. DiMER requires minimal setup. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mageeney, Catherine [Sandia National Lab. (SNL-CA)↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Pangenomes suggest ecological-evolutionary responses to experimental soil warming

ABSTRACT Below-ground carbon transformations that contribute to healthy soils represent a natural climate change mitigation, but newly acquired traits adaptive to climate stress may alter microbial feedback mechanisms. To better define microbial evolutionary responses to long-term climate warming, we study microorganisms from an ongoing in situ soil warming experiment where, for over three decades, temperate forest soils are continuously heated at 5°C above ambient. We hypothesize that across generations of chronic warming, genomic signatures within diverse bacterial lineages reflect adaptations related to growth and carbon utilization. From our bacterial culture collection isolated from experimental heated and control plots, we sequenced genomes representing dominant taxa sensitive to warming, including lineages of Actinobacteria, Alphaproteobacteria, and Betaproteobacteria. We investigated genomic attributes and functional gene content to identify signatures of adaptation. Comparative pangenomics revealed accessory gene clusters related to central metabolism, competition, and carbon substrate degradation, with few functional annotations explicitly associated with long-term warming. Trends in functional gene patterns suggest genomes from heated plots were relatively enriched in central carbohydrate and nitrogen metabolism pathways, while genomes from control plots were relatively enriched in amino acid and fatty acid metabolism pathways. We observed that genomes from heated plots had less codon bias, suggesting potential adaptive traits related to growth or growth efficiency. Codon usage bias varied for organisms with similar 16S rrn operon copy number, suggesting that these organisms experience different selective pressures on growth efficiency. Our work suggests the emergence of lineage-specific trends as well as common ecological-evolutionary microbial responses to climate change. IMPORTANCE Anthropogenic climate change threatens soil ecosystem health in part by altering below-ground carbon cycling carried out by microbes. Microbial evolutionary responses are often overshadowed by community-level ecological responses, but adaptive responses represent potential changes in traits and functional potential that may alter ecosystem function. We predict that microbes are adapting to climate change stressors like soil warming. To test this, we analyzed the genomes of bacteria from a soil warming experiment where soil plots have been experimentally heated 5°C above ambient for over 30 years. While genomic attributes were unchanged by long-term warming, we observed trends in functional gene content related to carbon and nitrogen usage and genomic indicators of growth efficiency. These responses may represent new parameters in how soil ecosystems feedback to the climate system.

Choudoir, Mallory J. (ORCID:0000000291175150)↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B↗

Opportunities and Challenges for Machine Learning-Assisted Enzyme Engineering

Enzymes can be engineered at the level of their amino acid sequences to optimize key properties such as expression, stability, substrate range, and catalytic efficiency or even to unlock new catalytic activities not found in nature. Because the search space of possible proteins is vast, enzyme engineering usually involves discovering an enzyme starting point that has some level of the desired activity followed by directed evolution to improve its “fitness” for a desired application. Recently, machine learning (ML) has emerged as a powerful tool to complement this empirical process. ML models can contribute to (1) starting point discovery by functional annotation of known protein sequences or generating novel protein sequences with desired functions and (2) navigating protein fitness landscapes for fitness optimization by learning mappings between protein sequences and their associated fitness values. In this Outlook, we explain how ML complements enzyme engineering and discuss its future potential to unlock improved engineering outcomes.

60 APPLIED LIFE SCIENCES↗

IMG Annotation Pipeline (IMGAP) v5.1.13

The IMG Annotation Pipeline is a collection of Bash and Python scripts to control a workflow for structural and functional annotation of prokaryotic genomes, metagenomes, and metatranscriptomes. The bash scripts in general control the overall workflow and are wrappers around 3rd party executables (not included in repo) that predict features or functions. Whereas the Python scripts do post-processing of raw output in terms of filtering or format transformation and in some cases contain some logic for picking the correct predictions or resolving overlaps. The pipeline is tailored to produce results required by IMG (https://img.jgi.doe.gov/) and is executed on every dataset submitted to IMG via https://img.jgi.doe.gov/submit. These consist of internal genomes, metagenomes and metatranscriptomes sequenced and assembled at the JGI, as well as datasets submitted by external users (non-lab/JGI affiliates).

Huntemann, Marcel↗

Reverse engineering environmental metatranscriptomes clarifies best practices for eukaryotic assembly

Abstract Background Diverse communities of microbial eukaryotes in the global ocean provide a variety of essential ecosystem services, from primary production and carbon flow through trophic transfer to cooperation via symbioses. Increasingly, these communities are being understood through the lens of omics tools, which enable high-throughput processing of diverse communities. Metatranscriptomics offers an understanding of near real-time gene expression in microbial eukaryotic communities, providing a window into community metabolic activity. Results Here we present a workflow for eukaryotic metatranscriptome assembly, and validate the ability of the pipeline to recapitulate real and manufactured eukaryotic community-level expression data. We also include an open-source tool for simulating environmental metatranscriptomes for testing and validation purposes. We reanalyze previously published metatranscriptomic datasets using our metatranscriptome analysis approach. Conclusion We determined that a multi-assembler approach improves eukaryotic metatranscriptome assembly based on recapitulated taxonomic and functional annotations from an in-silico mock community. The systematic validation of metatranscriptome assembly and annotation methods provided here is a necessary step to assess the fidelity of our community composition measurements and functional content assignments from eukaryotic metatranscriptomes.

Krinos, Arianna I. (ORCID:0000000197678392)↗

ETOP 503689: Pre-optimized cell free lysates for rapid prototyping of genes and pathways

This ETOP project aimed to re-conceive how we engineer complex biological systems by linking pathway design, prospecting, and validation into an integrated framework. Specifically, our vision seeks to advance and interweave high-throughput cell-based systems and rapid cell-free technologies in a way suitable for automation and microdroplet manipulation in a DOE JGI user facility setting. The key technology to be investigated toward this vision is a cell-free platform for combinatorial assembly of pathways by mixing-and-matching crude cell lysates derived from a suite of flux-enhanced background strains, each enriched with pathway enzymes. Through this work, we have established rewired E. coli and S. cerevisiae cells suitable to generated rewired lysates. We have demonstrated that lysates derived from rewired cells can generate higher fluxes of products and can enable a more robust discovery method going from gene to function annotations. Finally, we demonstrate that these rewired lysates can be stored for up to a year and still retain function. Collectively, these technologies enable a more rapid screening approach for enzyme and pathway variants.

59 BASIC BIOLOGICAL SCIENCES↗