Search NASASearch

SEARCH · Search NASA

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria

Shedding Light on Microbial Dark Matter with A Universal Language of Life

The majority of microbial genomes have yet to be cultured, and most proteins predicted from microbial genomes or sequenced from the environment cannot be functionally annotated. As a result, current computational approaches to describe microbial systems rely on incomplete reference databases that cannot adequately capture the full functional diversity of the microbial tree of life, limiting our ability to model high-level features of biological sequences. The scientific community needs a means to capture the functionally and evolutionarily relevant features underlying biology, independent of our incomplete reference databases. Such a model can form the basis for transfer learning tasks, enabling downstream applications in environmental microbiology, medicine, and bioengineering. Here we present LookingGlass, a deep learning model capturing a “universal language of life”. LookingGlass encodes contextually-aware, functionally and evolutionarily relevant representations of short DNA reads, distinguishing reads of disparate function, homology, and environmental origin. We demonstrate the ability of LookingGlass to be fine-tuned to perform a range of diverse tasks: to identify novel oxidoreductases, to predict enzyme optimal temperature, and to recognize the reading frames of DNA sequence fragments. LookingGlass is the first contextually-aware, general purpose pre-trained “biological language” representation model for short-read DNA sequences. LookingGlass enables functionally relevant representations of otherwise unknown and unannotated sequences, shedding light on the microbial dark matter that dominates life on Earth.

A Hoarfrost

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES

AI Foundation Models for Science: An Open Collaborative Initiative

Foundation Models (FMs), AI models designed to replace task-specific models, are increasingly being recognized for their versatility across numerous downstream applications. These models, trained using self-supervised techniques on any type of sequence data, circumvent the need for large annotated datasets, a major bottleneck in traditional AI model development. FMs can be applied to downstream tasks using few-shot learning and fine-tuning, significantly reducing the need for large labeled training datasets and computational resources. However, the development of FMs requires substantial resources, including access to data and compute power, expertise in the latest models, and specialized scientific knowledge for systematic evaluation. It is challenging for a single group to possess all these capabilities. To address this, NASA IMPACT has initiated an open collaborative effort, leveraging partnerships with the private sector and other groups within and outside NASA, to jointly build FMs. The overarching goal is to develop a consistent and collaborative approach to building FMs for high-value science datasets. This initiative has fostered collaboration within NASA and with external partners, including IBM Research, Clark University, DOE’s ORNL, ESA, and USGS. The effort focuses on identifying key datasets with a wide range of downstream applications, pretraining and building FMs using modified transformer architectures, evaluating compute infrastructure needs, and sharing models, pretraining and fine-tuning code, and data with the community. Furthermore, it aims to train the Earth science community to fine-tune these models for various downstream applications. Our initial effort resulted in the creation of a 100 million parameter HLS Geospatial Model within six months, which was released on HuggingFace. We are now expanding our scope to include data from weather and climate models and investigating multimodal models. We invite those interested in participating in this effort to join us by sharing their use cases, expertise, or data.

Rahul Ramachandran

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES

Telemetry-Enhancing Scripts

Scripts Providing a Cool Kit of Telemetry Enhancing Tools (SPACKLE) is a set of software tools that fill gaps in capabilities of other software used in processing downlinked data in the Mars Exploration Rovers (MER) flight and test-bed operations. SPACKLE tools have helped to accelerate the automatic processing and interpretation of MER mission data, enabling non-experts to understand and/or use MER query and data product command simulation software tools more effectively. SPACKLE has greatly accelerated some operations and provides new capabilities. The tools of SPACKLE are written, variously, in Perl or the C or C++ language. They perform a variety of search and shortcut functions that include the following: Generating text-only, Event Report-annotated, and Web-enhanced views of command sequences; Labeling integer enumerations with their symbolic meanings in text messages and engineering channels; Systematic detecting of corruption within data products; Generating text-only displays of data-product catalogs including downlink status; Validating and labeling of commands related to data products; Performing of convenient searches of detailed engineering data spanning multiple Martian solar days; Generating tables of initial conditions pertaining to engineering, health, and accountability data; Simplified construction and simulation of command sequences; and Fast time format conversions and sorting.

Maimone, Mark W.

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS

Analysis of genomic signatures associated with Variovorax endosphere colonization

This repository contains the analysis code and supporting datasets associated with the study “Genomic signatures in Variovorax enabling colonization of the Populus endosphere.” Beals DG, Carper DL, Hochanadel LH, Jawdy SS, Klingeman DM, Piatkowski BT, Weston DJ, Doktycz MJ, Pelletier DA. 2026. Genomic signatures in Variovorax enabling colonization of the Populus endosphere. mSystems 11:e01605-25. https://doi.org/10.1128/msystems.01605-25 The scripts are organized sequentially (01–07) and document the workflows used for: Sequence-read alignment and feature counting Orthogroup and KEGG Ortholog annotation Count normalization Statistical analysis and aggregation Generation of manuscript figures and tables Repository contents The uncompressed files are the finalized, formatted datasets used to generate the figures and tables reported in the study, including the supplemental CSV files referenced in the manuscript. The accompanying ZIP archive contains the complete codebase and example data_input/ and data_output/ directories illustrating the organization and execution of the analytical workflow. Individual scripts identify the corresponding manuscript analyses and figure panels. Raw sequencing data Raw sequencing reads are available through the NCBI Sequence Read Archive under BioProject accession PRJNA1322484.

Beals, Delaney [ORNL] (ORCID:0000000306274574)

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics

Survey of Thirteen Novel Pseudomonas putida Bacteriophages

Bacteriophages have been widely investigated as a promising treatment of food, medical equipment, and humans colonized by antibiotic-resistant bacteria. Phages pose particular interest in combating those bacteria which form biofilms, such as the medically important human pathogen Pseudomonas aeruginosa and several plant pathogens, including P. syringae . In an undergraduate lab course, P. putida was used as the host to isolate novel anti-pseudomonal bacteriophages. Environmental samples of soil and water were collected, and purified phage isolates were obtained. After Illumina sequencing, genomes of these phages were assembled de novo and annotated. Assembled genomes were compared with known genomes in the literature and GenBank to identify taxonomic relations and to refine their functional annotations. The thirteen phages described are sipho-, myo-, and podoviruses in several families of Caudoviricetes , spanning several novel genera, with genomes ranging from 40,000 to 96,000 bp. One phage (DDSR119) is unique and is the first reported P. putida siphovirus. The remaining 12 can be clustered into four distinct groups. Six are highly related to each other and to previously described Autotranscriptaviridae phages: Waldo5, PlaquesPlease, and Laces98 all belong to the Waldovirus genus, whereas Stalingrad, Bosely, and Stamos belong to the Troedvirus genus. Zuri was previously classified as the founding member of a new genus Zurivirus within the family Schitoviridae . Ebordelon and Holyagarpour each represent different species within Zurivirus , whereas Meara is a more distantly related member of the Schitoviridae . Dolphis and Jeremy are similar enough to form a genus but have only a few distant relatives among sequenced phages and are notable for being temperate. We identified the lysis cassettes in all 13 phages, compared tail spike structures, and found auxiliary metabolic genes in several. Studies like these, which isolate and characterize infectious virions, enable the identification of novel proteins and molecular systems and also provide the raw materials for further study, evaluation, and manipulation of phage proteins and their hosts.

Pseudomonas putida

A multi-omic characterization of the physiological responses to salt stress in Scenedesmus obliquus UTEX393

Scenedesmus obliquus UTEX393 is a promising microalgal candidate for sustainable biomanufacturing but its limited halotolerance hinders large-scale cultivation in saline environments. To investigate the molecular basis of salt stress responses, we conducted a comprehensive multi-omic analysis integrating genomics, transcriptomics, proteomics, lipidomics, metabolomics, and DNA affinity purification sequencing (DAP-seq). An improved nuclear genome assembly and annotation yielded 19,017 gene models and a 97% BUSCO completeness score, enabling construction of a genome-scale metabolic model. Comparing 15 ppt salinity stress to 5 ppt control, growth and productivity were significantly reduced, accompanied by widespread transcriptomic and proteomic changes. Transcriptomic analysis revealed downregulation of photosynthetic machinery and energy conservation genes, and upregulation of stress-responsive elements such as expansins, flavodoxins, and osmoprotectants. Lipidomic profiling showed accumulation of triacylglycerols (TAGs) and degradation of galactosyl lipids, consistent with a shift toward lipid biosynthesis to mitigate redox imbalance. Depletion of key polar metabolites and branched-chain amino acids suggested a rerouting of central carbon metabolism under stress. DAP-seq identified key transcription factors, including LHY1 and SPL12, that target central metabolic enzymes involved in redox balancing, such as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and malate dehydrogenase (MDH). These findings establish a regulatory-metabolic framework linking redox stress to lipid accumulation and reveal potential engineering targets to enhance salt tolerance. Overall, the multi-omic analysis supports the “overflow” hypothesis, where impaired photosynthesis results in excess reducing equivalents being diverted into TAG synthesis and highlights transcriptional regulators as candidates for improving algal robustness in brackish environments.

09 BIOMASS FUELS

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill

Rapid Diagnostics of Onboard Sequences

Keeping track of sequences onboard a spacecraft is challenging. When reviewing Event Verification Records (EVRs) of sequence executions on the Mars Exploration Rover (MER), operators often found themselves wondering which version of a named sequence the EVR corresponded to. The lack of this information drastically impacts the operators diagnostic capabilities as well as their situational awareness with respect to the commands the spacecraft has executed, since the EVRs do not provide argument values or explanatory comments. Having this information immediately available can be instrumental in diagnosing critical events and can significantly enhance the overall safety of the spacecraft. This software provides auditing capability that can eliminate that uncertainty while diagnosing critical conditions. Furthermore, the Restful interface provides a simple way for sequencing tools to automatically retrieve binary compiled sequence SCMFs (Space Command Message Files) on demand. It also enables developers to change the underlying database, while maintaining the same interface to the existing applications. The logging capabilities are also beneficial to operators when they are trying to recall how they solved a similar problem many days ago: this software enables automatic recovery of SCMF and RML (Robot Markup Language) sequence files directly from the command EVRs, eliminating the need for people to find and validate the corresponding sequences. To address the lack of auditing capability for sequences onboard a spacecraft during earlier missions, extensive logging support was added on the Mars Science Laboratory (MSL) sequencing server. This server is responsible for generating all MSL binary SCMFs from RML input sequences. The sequencing server logs every SCMF it generates into a MySQL database, as well as the high-level RML file and dictionary name inputs used to create the SCMF. The SCMF is then indexed by a hash value that is automatically included in all command EVRs by the onboard flight software. Second, both the binary SCMF result and the RML input file can be retrieved simply by specifying the hash to a Restful web interface. This interface enables command line tools as well as large sophisticated programs to download the SCMF and RMLs on-demand from the database, enabling a vast array of tools to be built on top of it. One such command line tool can retrieve and display RML files, or annotate a list of EVRs by interleaving them with the original sequence commands. This software has been integrated with the MSL sequencing pipeline where it will serve sequences useful in diagnostics, debugging, and situational awareness throughout the mission.

Starbird, Thomas W.

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES