Search NASASearch

SEARCH · Search NASA

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Energetics and Kinetics of Syntrophic Aromatic Degradation (Final Technical Report)

This DOE Basic Energy Physical Biosciences project spanned a 28-year period and two project investigators. Numerous significant research discoveries have occurred over this time frame. Much of this work has been reported in peer-reviewed literature, with the publication of results from the last four years forthcoming. The overarching project themes have focused on probing the metabolism of syntrophic bacteria and their environments. Research supported by this project has been central to ten doctoral dissertations and one master’s thesis at the University of Oklahoma. The project has also supported numerous undergraduate research projects throughout the years. Additionally, several microbial genomes were sequenced and annotated in association with this work 12-16. Key findings and publications from this work are highlighted. More detailed results are provided for unpublished and embargoed work. A complete list of publications and research products associated with this project is at the report's end.

59 BASIC BIOLOGICAL SCIENCES

Sensitive and error-tolerant annotation of protein-coding DNA with BATH

We present BATH, a tool for highly sensitive annotation of protein-coding DNA based on direct alignment of that DNA to a database of protein sequences or profile hidden Markov models (pHMMs). BATH is built on top of the HMMER3 code base, and simplifies the annotation workflow for pHMM-based translated sequence annotation by providing a straightforward input interface and easy-to-interpret output. BATH also introduces novel frameshift-aware algorithms to detect frameshift-inducing nucleotide insertions and deletions (indels). BATH matches the accuracy of HMMER3 for annotation of sequences containing no errors, and produces superior accuracy to all tested tools for annotation of sequences containing nucleotide indels. These results suggest that BATH should be used when high annotation sensitivity is required, particularly when frameshift errors are expected to interrupt protein-coding regions, as is true with long-read sequencing data and in the context of pseudogenes.

59 BASIC BIOLOGICAL SCIENCES

The Thiamine-Pyrophosphate-Motif

Thiamin pyrophosphate (TPP), a derivative of vitamin B1, is a cofactor for enzymes performing catalysis in pathways of energy production including the well known decarboxylation of a-keto acid dehydrogenases followed by transketolation. TPP-dependent enzymes constitute a structurally and functionally diverse group exhibiting multimeric subunit organization, multiple domains and two chemically equivalent catalytic centers. Annotation of functional TPP-dependcnt enzymes, therefore, has not been trivial due to low sequence similarity related to this complex organization. Our approach to analysis of structures of known TPP-dependent enzymes reveals for the first time features common to this group, which we have termed the TPP-motif. The TPP-motif consists of specific spatial arrangements of structural elements and their specific contacts to provide for a flip-flop, or alternate site, enzymatic mechanism of action. Analysis of structural elements entrained in the flip-flop action displayed by TPP-dependent enzymes reveals a novel definition of the common amino acid sequences. These sequences allow for annotation of TPP-dependent enzymes, thus advancing functional proteomics. Further details of three-dimensional structures of TPP-dependent enzymes will be discussed.

Ciszak, Ewa

Multi‐season analysis reveals hundreds of drought‐responsive genes in sorghum

Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.

Cole, Benjamin [USDOE Joint Genome Institute (JGI)

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii V.2

The head smut (Tilletia maclaganii) is a significant pathogen of the bioenergy crop switchgrass. T. maclaganii typically is more prevalent in older stands of switchgrass and can contribute to significant biomass loss. Here, we outline the methods for the sequencing, assembly, and annotation of the first reference genome for Tilletia maclaganii.

Benucci, Gian Maria Niccolò [GLBRC - Michigan Stat

GENEX: A knowledge-based expert assistant for Genbank data analysis

We describe a knowledge-based expert assistant, GENEX (Gene Explorer), that simplifies some analysis of Genbank data. GENEX is written in CLIPS (C Language Integrated Production System), and expert system tool, developed at the NASA Johnson Space Center. The main purpose of the system is to look for gene start site annotations, unusual DNA sequence composition, and regulatory protein patterns. application where determinations are made via a decision tree.

Batra, Sajeev

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)

Developing Dual Polarization Applications For 45th Weather Squadron's (45 WS) New Weather Radar: A Cooperative Project With The National Space Science and Technology Center (NSSTC)

A new weather radar is being acquired for use in support of America s space program at Cape Canaveral Air Force Station, NASA Kennedy Space Center, and Patrick AFB on the east coast of central Florida. This new radar includes dual polarization capability, which has not been available to 45 WS previously. The 45 WS has teamed with NSSTC with funding from NASA Marshall Spaceflight Flight Center to improve their use of this new dual polarization capability when it is implemented operationally. The project goals include developing a temperature profile adaptive scan strategy, developing training materials, and developing forecast techniques and tools using dual polarization products. The temperature profile adaptive scan strategy will provide the scan angles that provide the optimal compromise between volume scan rate, vertical resolution, phenomena detection, data quality, and reduced cone-of-silence for the 45 WS mission. The mission requirements include outstanding detection of low level boundaries for thunderstorm prediction, excellent vertical resolution in the atmosphere electrification layer between 0 C and -20 C for lightning forecasting and Lightning Launch Commit Criteria evaluation, good detection of anvil clouds for Lightning Launch Commit Criteria evaluation, reduced cone-of-silence, fast volume scans, and many samples per pulse for good data quality. The training materials will emphasize the appropriate applications most important to the 45 WS mission. These include forecasting the onset and cessation of lightning, forecasting convective winds, and hopefully the inference of electrical fields in clouds. The training materials will focus on annotated radar imagery based on products available to the 45 WS. Other examples will include time sequenced radar products without annotation to simulate radar operations. This will reinforce the forecast concepts and also allow testing of the forecasters. The new dual polarization techniques and tools will focus on the appropriate applications for the 45 WS mission. These include forecasting the onset of lightning, the cessation of lightning, convective winds, and hopefully the inference of electrical fields in clouds. This presentation will report on the results achieved so far in the project.

Roeder, W.P.

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES

A genomic perspective on fungal diversity and evolution

Originating from aquatic unicellular ancestors, over the course of ~1 billion years, the fungi have evolved to occupy nearly all aerobic environments on the planet, diversified into millions of different ‘species’ and have developed complex multicellular structures. Their relatively small, simple genomes have facilitated massive-scale sequencing and allowed us to explore genome evolution across an ancient eukaryotic kingdom. With thousands of genomes from diverse lineages now available, this Review will discuss insights into fungal biology and evolution gleaned with genomics and other multi-omics approaches. Using published genomes available through GenBank and the Joint Genome Institute’s MycoCosm platform, we generated kingdom-wide phylogenies and used them to highlight how fungal genomes have changed over time. With this phylogeny as a guide, we also discuss major evolutionary transitions that occurred across the fungal kingdom. Although progress has been made, these efforts are hampered by biases in genome representation and limited characterization of gene functions. Here, in this study, we discuss these challenges and possible future directions to address them, including initiatives to characterize conserved genes of unknown function and scale up sequencing towards 10,000 annotated fungal genomes.

Mondo, Stephen J. [USDOE Joint Genome Institute (J

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto

Cleanroom Microbes Survive Drying, Vacuum, and Proton Irradiation

Introduction : The goal of planetary protection at NASA is to mitigate the risk of contaminating sensitive target bodies with biological life. While many cleaning procedures have been put in place to reduce bioburden on spacecraft, microbes are experts at evolving to survive harsh conditions. Specifically, the dry, low-nutrient environment of a cleanroom (commonly used for assembly of spacecraft) can represent an environment where extremophiles can survive. Methods : Scientists at NASA MSFC wished to gather a snapshot of the microbial population within a variety of cleanrooms on site. A study was undertaken to collect air, surface, and floor samples from clean-rooms and isolate unique morphologies. From this study, 95 isolates were collected and saved in a microbial library. About 86% of these were identified at least to a genus level. Following identification, 24 microbes were selected, based on a literature review, as potential extremophiles. These were grown in liquid cultures, diluted to a set optical density, washed with water, and then applied to a sterilized Kapton coupon. Droplets were allowed to dry overnight in a biosafety cabinet. Coupons were then installed in a pelletron and pumped down to high vacuum (~1E-6 Torr). Samples were then subjected 100 keV protons at a fluence of 2x10 15 p+/cm 2 up to 4x10 15 p+/cm 2 . Following exposure, samples were returned to the microbiology lab where they were pro-cessed by submerging in water, vortexing, and then plating either droplets or spread plates. Recovery data collected was qualitative with a ranking or +, minor, or – for growth. Some selected radiotolerant strains were sequenced using the Illumina sequencing platform. The resulting genomes were annotated with the Rapid Annotations using Subsystems Technology (RAST) server and analyzed for conserved and unique stress response relevant genomic signatures to identify clues related to specific tolerances. Results and Discussion : After five rounds of proton radiation, we narrowed our isolates to five, non-spore forming bacteria that demonstrated survival: Arthrobacter koreensis, Paenarthrobacter nitroguajacolicus, Mycetocola manganoxydans , and an Erwinia sp. Furthermore, we exposed these four microbes to 254 nm wavelength light at an intensity of 80 W/m 2 at a distance of ~18 cm for 10 minutes. Only A. koreensis demonstrated survival following UV exposure. Finally, we performed whole genome sequencing on the four strains to look for genetic markers of stress resistance. When we compared the genomes of the four strains, we found that genes coding for GGDEF and EAL domains with PAS/PAC sensors were only found in A. koreensis . These domains, modulated by PAS/PAC sensors, are hypothesized to facilitate survival under drying, desiccation, and proton irradiation. Drying and Desiccation : PAS domains sense hydration changes and modulate GGDEF and EAL domain activity to adjust c-di-GMP levels, enhancing resistance to desiccation. For instance, in Pseudomonas aeruginosa , the PAS domain of RbdA modulates activity under varying hydration conditions, affecting stress responses [1]. Proton Irradiation : Proton irradiation causes oxidative stress, leading to ROS generation. PAS domains detect this stress and modulate GGDEF and EAL domains to manage oxidative stress responses. In Shewanella , EAL domain proteins modulated by PAS sensors help bacteria adapt to extreme conditions [2]. These genes upregulate other stress response genes, protecting membrane function, protein stability, DNA repair, and antioxidant defenses. The modulation of c-di-GMP by PAS domains is crucial for bacterial adaptation to stress conditions, enabling dynamic physio-logical adjustments [3]. Understanding these mechanisms provides insights into bacterial stress responses and strategies for controlling bacterial growth [4]. Conclusions : These findings indicate that clean-rooms harbor extremophile microbes that may be able to survive conditions in deep space. Furthermore, while we identified certain stress-response genes that may be at least partly responsible for the phenotypes observed in this study, there are likely unidentified genes or characteristics about A. koreensis , and other bacteria, that may allow them to survive in harsh environments. Future studies will focus on identifying these unknown genes and characteristics, further elucidating the mechanisms of extremophile survival and potentially informing the development of new biotechnologies for space exploration and other extreme environments.

Chelsi Cassilly