Search NASASearch

SEARCH · Search NASA

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor

Local frustration around enzyme active sites

Conflicting biological goals often meet in the specification of protein sequences for structure and function. Overall, strong energetic conflicts are minimized in folded native states according to the principle of minimal frustration, so that a sequence can spontaneously fold, but local violations of this principle open up the possibility to encode the complex energy landscapes that are required for active biological functions. We survey the local energetic frustration patterns of all protein enzymes with known structures and experimentally annotated catalytic residues. In agreement with previous hypotheses, the catalytic sites themselves are often highly frustrated regardless of the protein oligomeric state, overall topology, and enzymatic class. At the same time a secondary shell of more weakly frustrated interactions surrounds the catalytic site itself. We evaluate the conservation of these energetic signatures in various family members of major enzyme classes, showing that local frustration is evolutionarily more conserved than the primary structure itself.

Maria I. Freiberger

Fungal diversity and function in metagenomes sequenced from extreme environments

Fungi are increasingly recognized as key players in various extreme environments. Here we present an analysis of publicly-sourced metagenomes from global extreme environments, focusing on fungal taxonomy and function. The majority of 855 selected metagenomes contained scaffolds assigned to fungi. Relative abundance of fungi was as high as 10% of protein-coding genes with taxonomic annotation, with up to 289 fungal genera per sample. Despite taxonomic clustering by environment, fungal communities were more dissimilar than archaeal and bacterial communities, both for within- and between-environment comparisons. Relatively abundant fungal classes in extreme environments included Dothideomycetes, Eurotiomycetes, Leotiomycetes, Pezizomycetes, Saccharomycetes, and Sordariomycetes. Broad generalists and prolific aerial spore formers were the most relatively abundant fungal genera detected in most of the extreme environments, bringing up the question of whether they are actively growing in those environments or just surviving as spores. More specialized fungi were common in some environments, such as zoosporic taxa in cryosphere water and hot springs. Relative abundances of genes involved in adaptation to general, thermal, oxidative, and osmotic stress were greatest in soda lake, acid mine drainage, and cryosphere water samples.

60 APPLIED LIFE SCIENCES

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

Pangenomes suggest ecological-evolutionary responses to experimental soil warming

ABSTRACT Below-ground carbon transformations that contribute to healthy soils represent a natural climate change mitigation, but newly acquired traits adaptive to climate stress may alter microbial feedback mechanisms. To better define microbial evolutionary responses to long-term climate warming, we study microorganisms from an ongoing in situ soil warming experiment where, for over three decades, temperate forest soils are continuously heated at 5°C above ambient. We hypothesize that across generations of chronic warming, genomic signatures within diverse bacterial lineages reflect adaptations related to growth and carbon utilization. From our bacterial culture collection isolated from experimental heated and control plots, we sequenced genomes representing dominant taxa sensitive to warming, including lineages of Actinobacteria, Alphaproteobacteria, and Betaproteobacteria. We investigated genomic attributes and functional gene content to identify signatures of adaptation. Comparative pangenomics revealed accessory gene clusters related to central metabolism, competition, and carbon substrate degradation, with few functional annotations explicitly associated with long-term warming. Trends in functional gene patterns suggest genomes from heated plots were relatively enriched in central carbohydrate and nitrogen metabolism pathways, while genomes from control plots were relatively enriched in amino acid and fatty acid metabolism pathways. We observed that genomes from heated plots had less codon bias, suggesting potential adaptive traits related to growth or growth efficiency. Codon usage bias varied for organisms with similar 16S rrn operon copy number, suggesting that these organisms experience different selective pressures on growth efficiency. Our work suggests the emergence of lineage-specific trends as well as common ecological-evolutionary microbial responses to climate change. IMPORTANCE Anthropogenic climate change threatens soil ecosystem health in part by altering below-ground carbon cycling carried out by microbes. Microbial evolutionary responses are often overshadowed by community-level ecological responses, but adaptive responses represent potential changes in traits and functional potential that may alter ecosystem function. We predict that microbes are adapting to climate change stressors like soil warming. To test this, we analyzed the genomes of bacteria from a soil warming experiment where soil plots have been experimentally heated 5°C above ambient for over 30 years. While genomic attributes were unchanged by long-term warming, we observed trends in functional gene content related to carbon and nitrogen usage and genomic indicators of growth efficiency. These responses may represent new parameters in how soil ecosystems feedback to the climate system.

Choudoir, Mallory J. (ORCID:0000000291175150)

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits and two catalytic centers. Each catalytic center (PP:PYR) is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and amhopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core (PP:PYR)(sub 2) within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GXPhiX(sub 4)(G)PhiXXGQ and GDGX(sub 25-30)NN in the PP-domain, and the EX(sub 4)(G)PhiXXGPhi in the PYR-domain, where Phi corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, P.

LibraryX: A Framework for Cross-Library-Call Optimization

Scientific applications utilize performance libraries as a software engineering concept: these libraries encapsulate important and well-understood (mathematical) operations, allow for reuse, and are implemented and tuned by experts. Domain scientists then implement complex algorithms based on these domainspecific libraries. While individual library calls are optimized, larger performance gains across sequences of calls—sometimes spanning multiple libraries—are often unrealized, forcing a trade-off between performance and implementation complexity.To overcome this issue, we propose LibraryX, an approach and a system that allows for cross-library-call optimization even when library calls stem from multiple performance libraries. LibraryX annotates library calls with semantic information and optimizes entire directed acyclic graphs (DAGs) of calls dynamically using the SPIRAL code generation system. We demonstrate its effectiveness across a range of memory bound workloads, achieving significant speedups on Nvidia, AMD, and Intel accelerators compared to code using native libraries without cross-call optimization.

Rao, Sanil [Carnegie Mellon University,Department

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

Verification of Numerical Algorithms

The following strategy is suggested for specification and proof: (1) Defer the construction of a formal program specification with respect to I/O assertions unit the correctness of the program with respect to an abstract mathematical model of program intent is demonstrated. (2) Prove that an abstract machine (using infinite precision arithmetic) would compute that object exactly. (3) Prove that the computational sequences of arithmetic operations that occur in the abstract machine must be precisely the same at every step as those occurring on an actual machine (with finite precision arithmetic), executing the same program. (4) Use a Verification Conditions VC-generator that knows about the semantics of arithmetic operations to annotate the program with assertions that bound (or in some circumstances estimate) the difference between the actual machine state variables and the corresponding ones of the abstract machine. Construct the formal program specification by combining the verification conditions into theorems about computational error that can be proved with mechanical assistance.

Source record

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES

Tissue Photolithography

Tissue lithography will enable physicians and researchers to obtain macromolecules with high purity (greater than 90 percent) from desired cells in conventionally processed, clinical tissues by simply annotating the desired cells on a computer screen. After identifying the desired cells, a suitable lithography mask will be generated to protect the contents of the desired cells while allowing destruction of all undesired cells by irradiation with ultraviolet light. The DNA from the protected cells can be used in a number of downstream applications including DNA sequencing. The purity (i.e., macromolecules isolated form specific cell types) of such specimens will greatly enhance the value and information of downstream applications. In this method, the specific cells are isolated on a microscope slide using photolithography, which will be faster, more specific, and less expensive than current methods. It relies on the fact that many biological molecules such as DNA are photosensitive and can be destroyed by ultraviolet irradiation. Therefore, it is possible to protect the contents of desired cells, yet destroy undesired cells. This approach leverages the technologies of the microelectronics industry, which can make features smaller than 1 micrometer with photolithography. A variety of ways has been created to achieve identification of the desired cell, and also to designate the other cells for destruction. This can be accomplished through chrome masks, direct laser writing, and also active masking using dynamic arrays. Image recognition is envisioned as one method for identifying cell nuclei and cell membranes. The pathologist can identify the cells of interest using a microscopic computerized image of the slide, and appropriate custom software. In one of the approaches described in this work, the software converts the selection into a digital mask that can be fed into a direct laser writer, e.g. the Heidelberg DWL66. Such a machine uses a metalized glass plate (with chrome metallization) on which there is a thin layer of photoresist. The laser transfers the digital mask onto the photoresist by direct writing, with typical best resolution of 2 micrometers. The plate is then developed to remove the exposed photoresist, which leaves the exposed areas susceptible to chemical chrome etch. The etch removes the unprotected chrome. The rest of the photoresist is then removed, by either ultraviolet organic solvent or over-development. The remaining chrome pattern is quickly oxidized by atmospheric exposure (typically within 30 seconds). The ready chrome mask is now applied to the tissue slide and aligned manually, or using automatic software and pre-designed alignment marks. The slide plate sandwich is then exposed to UV to destroy the DNA of the unwanted cells. The slide and plate are separated and the slide is processed in a standard way to prepare for polymerase chain reaction (PCR) and potential identification of cancer sequences.

Wade, Lawrence A.

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO

Barcoded overexpression screens in gut Bacteroidales identify genes with roles in carbon utilization and stress resistance

Abstract A mechanistic understanding of host-microbe interactions in the gut microbiome is hindered by poorly annotated bacterial genomes. While functional genomics can generate large gene-to-phenotype datasets to accelerate functional discovery, their applications to study gut anaerobes have been limited. For instance, most gain-of-function screens of gut-derived genes have been performed in Escherichia coli and assayed in a small number of conditions. To address these challenges, we develop Barcoded Overexpression BActerial shotgun library sequencing (Boba-seq). We demonstrate the power of this approach by assaying genes from diverse gut Bacteroidales overexpressed in Bacteroides thetaiotaomicron . From hundreds of experiments, we identify new functions and phenotypes for 29 genes important for carbohydrate metabolism or tolerance to antibiotics or bile salts. Highlights include the discovery of a d -glucosamine kinase, a raffinose transporter, and several routes that increase tolerance to ceftriaxone and bile salts through lipid biosynthesis. This approach can be readily applied to develop screens in other strains and additional phenotypic assays.

59 BASIC BIOLOGICAL SCIENCES

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics