Search NASA⌕ Search

SEARCH · Search NASA

Results for “annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenome-assembled-genomes recovered from the Arctic drift expedition MOSAiC

The Multidisciplinary Observatory for Study of the Arctic Climate (MOSAiC) expedition consisted of a year-long drifting survey of the Central Arctic Ocean. The ecosystems component of MOSAiC included the sampling of molecular data, with metagenomes collected from a diverse range of environments. The generation of metagenome-assembled-genomes (MAGs) from metagenomes are a starting point for genome-resolved analyses. This dataset presents a catalogue of MAGs recovered from a set of 73 samples from MOSAiC, including 2407 prokaryotic and 56 eukaryotic MAGs, as well as annotations of a near complete eukaryotic MAG using the Joint Genome Institute (JGI) annotation pipeline. The metagenomic samples are from the surface ocean, chlorophyll maximum, mesopelagic and bathypelagic, within leads and under-ice ocean, as well as melt ponds, ice ridges, and first- and second-year sea ice. This set of MAGs can be used to benchmark microbial biodiversity in the Central Arctic Ocean, compare individual strains across space and time, and to study changes in Arctic microbial communities from the winter to summer, at a genomic level.

59 BASIC BIOLOGICAL SCIENCES↗

Patch-Based Convolutional Neural Networks for Multiple Microstructural Features Detection in FIB-SEM Micrographs of Irradiated Nuclear Fuel

Focused ion beam scanning electron microscopy (FIB-SEM) tomography has increasingly been utilized for acquiring three-dimensional (3D) microstructure features at the sub-micron scale in irradiated nuclear materials. This technique involves sequential ion beam slicing followed by electron beam imaging and compositional mapping using energy dispersive spectroscopy (EDS). Despite its growing use, several challenges persist. These include the time-intensive nature of data collection of EDS data, difficulties in distinguishing between various microstructures, and issues with image alignment. These challenges currently limit the broader application of FIB-SEM tomography in the field. To overcome these limitations, we propose using convolutional neural networks (CNNs) to automate microstructure identification in SEM images. Our study introduces a new framework for identifying microstructures in irradiated U-10Zr (wt. %) metallic fuel with limited annotated data. The framework includes the creation of a reliable annotated dataset with paired SEM and ground truth data from EDS maps, the applications of CNNs for microstructure identification, and the validation of model performance. Specifically, we employed the Segment Anything Model (SAM) to align SEM images with corresponding EDS maps and focused ion beam (FIB) tomography SEM data. We evaluate several models, including Patch-based U-Net, Attention U-Net, and Residual U-Net, finding that patch-based U-Net exhibits superior segmentation performance and consistency. This approach reduces reliance on EDS detectors and aids in accelerating nuclear material analysis process, highlighting the potential of advanced deep learning techniques to improve microstructural understanding in nuclear material. This is the first framework to integrate SAM and Patch-based CNN models for semantic segmentation of irradiated nuclear materials, with potential applicability to other tomography datasets.

36 - MATERIALS SCIENCE↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

Relics of interspecific hybridization retained in the genome of a drought-adapted peanut cultivar

Peanut (Arachis hypogaea L.) is a globally important oil and food crop frequently grown in arid, semi-arid, or dryland environments. Improving drought tolerance is a key goal for peanut crop improvement efforts. Here, we present the genome assembly and gene model annotation for “Line8,” a peanut genotype bred from drought-tolerant cultivars. Our assembly and annotation are the most contiguous and complete peanut genome resources currently available. The high contiguity of the Line8 assembly allowed us to explore structural variation both between peanut genotypes and subgenomes. We detect several large inversions between Line8 and other peanut genome assemblies, and there is a trend for the inversions between more genetically diverged genotypes to have higher gene content. We also relate patterns of subgenome exchange to structural variation between Line8 homeologous chromosomes. Unexpectedly, we discover that Line8 harbors an introgression from A.cardenasii, a diploid peanut relative and important donor of disease resistance alleles to peanut breeding populations. The fully resolved sequences of both haplotypes in this introgression provide the first in situ characterization of A.cardenasii candidate alleles that can be leveraged for future targeted improvement efforts. The completeness of our genome will support peanut biotechnology and broader research into the evolution of hybridization and polyploidy.

60 APPLIED LIFE SCIENCES↗

A haplotype-resolved reference genome for Eucalyptus grandis

Eucalyptus grandis is a hardwood tree used worldwide as pure species or hybrid partner to breed fast-growing plantation forestry crops that serve as feedstocks of timber and lignocellulosic biomass for pulp, paper, biomaterials, and biorefinery products. The current v2.0 genome reference for the species served as the first reference for the genus and has helped drive the development of molecular breeding tools for eucalypts. Using PacBio HiFi long reads and Omni-C proximity ligation sequencing, we produced an improved, haplotype-phased assembly (v4.0) for TAG0014, an early-generation selection of E. grandis. The 2 haplotypes are 571 Mbp (HAP1) and 552 Mbp (HAP2) in size and consist of 37 and 46 contigs scaffolded onto 11 chromosomes (contig N50 of 28.9 and 16.7 Mbp), respectively. These haplotype assemblies are 70-90 Mbp smaller than the diploid v2.0 assembly but capture all except one of the 22 telomeres, suggesting that substantial redundant sequence was included in the previous assembly. A total of 35,929 (HAP1) and 35,583 (HAP2) gene models were annotated, of which 438 and 472 contain long introns (>10 kbp) in gene models previously (v2.0) identified as multiple smaller genes. These and other improvements have increased gene annotation completeness levels from 93.8 to 99.4% in the v4.0 assembly. We found that 6,493 and 6,346 genes are within tandem duplicate arrays (HAP1 and HAP2, respectively, 18.4 and 17.8% of the total) and >43.8% of the haplotype assemblies consists of repeat elements. Analysis of synteny between the haplotypes and the E. grandis v2.0 reference genome revealed extensive regions of collinearity, but also some major rearrangements, and provided a preview of population and pangenome variation in the species.

Lötter, Anneri↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

The Gene Ontology knowledgebase in 2026

Abstract The Gene Ontology (GO) knowledgebase (https://geneontology.org) is a comprehensive resource describing the functions of genes. The GO knowledgebase is regularly updated and improved. We describe here the major updates that have been made in the past 3 years. The ontology and annotations have been expanded and revised, particularly in several areas of biology: cellular metabolism, multi-organism interactions (e.g. host-pathogen), extracellular matrix proteins, chromatin remodeling (e.g. the “histone code”), and noncoding RNA functions. We have released version 2 of a comprehensive set of integrated, reviewed annotations for human genes, which we call the “functionome.” We have also dramatically increased the number of GO-CAM models, with over 1500 models of metabolic and signaling pathways, primarily in human, mouse, budding and fission yeast, and fruit fly. Finally, we discuss our current recommendations and future prospects of AI in the use and development of GO.

Aleksander, Suzi A (ORCID:0000000167872901)↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

Signature analysis of high-throughput transcriptomics screening data for mechanistic inference and chemical grouping

Abstract High-throughput transcriptomics (HTTr) uses gene expression profiling to characterize the biological activity of chemicals in in vitro cell-based test systems. As an extension of a previous study testing 44 chemicals, HTTr was used to screen an additional 1,751 unique chemicals from the EPA’s ToxCast collection in MCF7 cells using 8 concentrations and an exposure duration of 6 h. We hypothesized that concentration-response modeling of signature scores could be used to identify putative molecular targets and cluster chemicals with similar bioactivity. Clustering and enrichment analyses were conducted based on signature catalog annotations and ToxPrint chemotypes to facilitate molecular target prediction and grouping of chemicals with similar bioactivity profiles. Enrichment analysis based on signature catalog annotation identified known mechanisms of action (MeOAs) associated with well-studied chemicals and generated putative MeOAs for other active chemicals. Chemicals with predicted MeOAs included those targeting estrogen receptor (ER), glucocorticoid receptor (GR), retinoic acid receptor (RAR), the NRF2/KEAP/ARE pathway, AP-1 activation, and others. Using reference chemicals for ER modulation, the study demonstrated that HTTr in MCF7 cells was able to stratify chemicals in terms of agonist potency, distinguish ER agonists from antagonists, and cluster chemicals with similar activities as predicted by the ToxCast ER Pathway model. Uniform manifold approximation and projection (UMAP) embedding of signature-level results identified novel ER modulators with no ToxCast ER Pathway model predictions. Finally, UMAP combined with ToxPrint chemotype enrichment was used to explore the biological activity of structurally related chemicals. The study demonstrates that HTTr can be used to inform chemical risk assessment by determining in vitro points of departure, predicting chemicals’ MeOA and grouping chemicals with similar bioactivity profiles.

Toxicology↗

DeepAndes: A Self-Supervised Vision Foundation Model for Multispectral Remote Sensing Imagery of the Andes

By mapping sites at large scales usingremotely sensed data, archaeologists can generate unique insights into long-term demographic trends, interregional social networks, and human adaptations in the past. Remote sensing surveys complement field-based approaches, and their reach can be especially great when combined with deep learning and computer vision techniques. However, conventional supervised deep learning methods face challenges in annotating fine-grained archaeological features at scale. In addition, while recent vision foundation models have shown remarkable success in learning large-scale remote sensing data with minimal annotations, most off-the-shelf solutions are designed for RGB images rather than multispectral satellite imagery, such as the eight-band data used in our study. In this article, we introduce DeepAndes, a transformer-based vision foundation model trained on three million multispectral satellite images, specifically tailored for Andean archaeology. DeepAndes incorporates a customized DINOv2 self-supervised learning algorithm optimized for eight-band multispectral imagery, marking the first foundation model designed explicitly for the Andes region. We evaluate its image understanding performance through imbalanced image classification, image instance retrieval, and pixel-level semantic segmentation tasks. Our experiments show that DeepAndes achieves superior F1 scores, mean average precision, and Dice scores in few-shot learning scenarios, significantly outperforming models trained from scratch or pretrained on smaller datasets. This underscores the effectiveness of large-scale self-supervised pretraining in archaeological remote sensing.

Guo, Junlin [Vanderbilt Univ., Nashville, TN (Unit↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

BiGEST

Natural products have provided a rich reservoir of beneficial compounds in public health including antibiotics, therapeutics, and immunosuppressants. These natural products are synthesized by enzymes encoded by Biosynthetic Gene Clusters (BGCs), clusters of co-localized biosynthetic genes. Computational detection of BGCs has become a crucial step in natural product discovery. While this process has been facilitated in bacterial and fungal organisms thanks to the currently available tools (e.g., antiSMASH), a large spectrum of eukaryotic organisms have been neglected by these existing tools due to the scarcity and incompleteness of genome annotation resources. Here, we introduce Biosynthetic Gene cluster Extensive Search Tool (BiGEST) to provide an extensive annotation-free search for BGCs in diverse eukaryotic organisms. As a result, BiGEST uncovers eukaryotic BGCs that could be undetected by other BGC detection tools.

Adriani, Lisa↗

DiMER

SAND2025-04145O DiMER is a Python based tool that helps researchers understand the functions of genes by searching through multiple biological databases. It takes user-provided data and scans various databases to find the best matches for gene functions, generating a clear summary of results. DiMER identifies the most relevant functional annotations and improves upon previous annotations by replacing instances of "unknown protein function" with more accurate descriptions. DiMER requires minimal setup. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mageeney, Catherine [Sandia National Lab. (SNL-CA)↗

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗