Search NASA⌕ Search

SEARCH · Search NASA

Results for “annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Inventory of Composable Elements (ICE) v6.0.0

The Inventory of Composable Elements (ICE) is an open source registry software platform for managing information about biological parts. It is capable of recording information about plasmids, microbial host strains and seeds, as well as DNA parts. Includes features such as DNA sequence visualization, editing and annotation, auto-aligning sequencing trace files against reference templates, SBOL XML/RDF support, and web-of-registries functionality. The web of registries functionality provides strong support for distributed interconnected use and enables sharing and transfer of biological parts across various independent ICE instances. ICE adopts modern software development principles, leveraging component-base frameworks, offering a REST API for convenient third-party integration and emphasizing scalability, security, and service integrations for dynamic content availability. The source code is hosted at https://github.com/JBEI/ice. A public instance is available at public-registry.jbei.org, where users can try out features, upload parts or simply use it for their projects.

Plahar, Hector↗

Biological Parts Search Portal (BioParts) v1.0.0

BioParts is a web based search portal for biological parts available in the public domain. It combines the ease and convenience of modern web search engines with the capabilities of bioinformatics search tools such as BLAST. This portal, available at bioparts.org, allows anyone to search for publicly accessible biological part information (e.g., NCBI, iGEM, SynBioHub, Addgene), including parts publicly accessible through ICE Registries. Additionally, the portal offers a REST API that enables third-party applications and tools to access the portal's functionality programmatically. While there are several standalone biological part repositories, there doesn't exist an application that indexes these publicly available parts and enables features such as keyword and BLAST searches along with automatic sequence annotation.

Plahar, Hector↗

Post Irradiation Examination Dislocation Defect Detection Software

This software provides dislocation-type defect identification and segmentation using a standard open source computer vision model, YOLOv8, that leverages transfer learning to create a highly effective dislocation defect quantification tool while using only a minimal number of expert annotated micrographs for training. This model demonstrates the ability to segment both dislocation lines and loops concurrently in micrographs with high pixel noise levels and on multiple alloys. It includes multiple layers of frozen layers used for transfer learning from multidisciplinary data and is extensible to alloys that are not included in the training dataset.

Anderson, MatthewW↗

pnnl-predictive-phenomics/csc052cyc

Using the genome annotation as input, Pathway-tools generates a database containing all the information that can be inferred from the genome. The Pathway/Genome database (PGDB) can subsequently be curated manually Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

pnnl-predictive-phenomics/csc040cyc

Using the genome annotation as input, Pathway-tools generates a database containing all the information that can be inferred from the genome. The Pathway/Genome database (PGDB) can subsequently be curated manually

Zucker, Jeremy [Pacific Northwest National Laborat↗

pnnl-predictive-phenomics/csc009cyc

Using the genome annotation as input, Pathway-tools generates a database containing all the information that can be inferred from the genome. The Pathway/Genome database (PGDB) can subsequently be curated manually. Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

SF-25-047 IMAGEMARKER

Image Marker is a tool for marking, categorizing, and annotating TIFF, FITS, PNG, and JPEG files. The software is intended to facilitate crowd sourcing human classification of features in scientific images without having to use web tools.

BLEEM, LINDSEYE [Argonne National Laboratory (ANL)↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

Radsource Mr: Mixed Reality Planning Tool For Radioactive Recovery

The RadSource MR system leverages Meta Quest 3's advanced mixed reality capabilities to create a comprehensive spatial planning platform for end-of-life sealed radioactive source recovery operations. The application utilizes the Quest 3's high-resolution passthrough cameras and spatial mapping algorithms to generate accurate 3D environmental models. Core technical components include: (1) Real-time spatial measurement algorithms calculating distances, angles, slopes, and surface areas with sub-centimeter accuracy; (2) Virtual object placement system allowing users to position digital representations of recovery equipment (trailers, containment vessels, protective barriers) within the real environment; (3) Voice recording and annotation system for hands-free documentation in protective equipment; (4) 3D mesh capture and storage capabilities for post-operation analysis and regulatory documentation. (5) Procedure documentation is available for viewing in Mixed Reality, providing an innovative and convenient way to access the information during pre-visit and inspection activities. (6) Support for screen capture for the view for real world and virtual objects together to use it later for planning. The system integrates computer vision techniques for environmental understanding, spatial mathematics for precise measurements, and human-computer interaction principles optimized for hazardous environment operations. Data persistence allows teams to save and share planning sessions across multiple stakeholders while maintaining operational security requirements.

Khadka, Rajiv [Idaho National Laboratory (INL), Id↗

CodeScribe Agent

SF-26-086 CodeScribe introduces a structured, multi-stage pipeline that combines deterministic program analysis with LLM-powered translation to enable incremental, testable Fortran-to-C++ migration. First, `code-scribe index` traverses the project directory tree and produces `scribe.yaml` metadata files recording all modules, subroutines, and functions at each level, giving the LLM accurate structural context instead of a hallucinated codebase model. Second, `code-scribe draft` performs the deterministic portion of translation — converting Fortran types to C++ equivalents, replacing `use` statements with `#include` and `using namespace` directives, and detecting constructs requiring special handling — while embedding`scribe-prompt` annotations that guide the LLM through non-trivial cases such as statement-function-to-lambda conversions and `extern "C"` wrapper generation. Third, `code-scribe translate` applies project-specific TOML-based few-shot prompt templates and submits the composed prompt to a pluggable LLM backend (OpenAI, Anthropic, Argonne ARGO, any OpenAI-compatible endpoint, or local Hugging Face checkpoints), producing a C++ source file, a header, and a Fortran-C++ interface file for each translated routine so the codebase compiles and runs correctly throughout the migration. Beyond translation, CodeScribe includes a tool-using coding agent (`code-scribe agent`) with read, bash, edit, and write capabilities, and a bounded loop mode (`code-scribe loop`) that runs repeated stateless agent sessions over a task file with restricted tool access — enabling sustained, auditable software development workflows for broader scientific computing tasks.

Dhruv, Akash [Argonne National Laboratory (ANL), A↗

Pyctos

Pyctos is a concolic testing framework for standard, dynamically-typed Python. Pyctos is able to generate exhaustive test inputs reaching 100% coverage for a subset of pure Python in the absence of type annotations, even where modern fuzzers would fail. Pyctos's underlying reasoning engine is the CVC5 SMT solver, though Z3 is also supported. Pyctos also supports a growing subset of the standard library and some third-party libraries, such as NumPy.

Washbourne, ErickN [Lawrence Livermore National La↗

ClusterWeave

ClusterWeave is a workflow for biosynthetic target discovery and prioritization. It assembles annotation, BGC detection, BiG-SCAPE family context, shortlist generation, and clinker-ready panel staging into one reproducible workflow.

Martin, StantonL↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

The Artificial Intelligence Ontology: LLM-Assisted Construction of AI Concept Hierarchies

The Artificial Intelligence Ontology (AIO) is a systematization of artificial intelligence (AI) concepts, methodologies, and their interrelations. Developed via manual curation, with the additional assistance of large language models (LLMs), AIO aims to address the rapidly evolving landscape of AI by providing a comprehensive framework that encompasses both technical and ethical aspects of AI technologies. The primary audience for AIO includes AI researchers, developers, and educators seeking standardized terminology and concepts within the AI domain. We use the term “branches” for classes, and their subclasses, in our ontology that are subclasses of owl:Thing. AIO contains eight branches: Bias, Layer, Machine Learning Task, Mathematical Function, Model, Network, Preprocessing, and Training Strategy, each designed to support the modular composition of AI methods and facilitate a deeper understanding of deep learning architectures and ethical considerations in AI. AIO uses the Ontology Development Kit (ODK) for its creation and maintenance, with its content being more easily updated through AI-driven curation support. This approach not only ensures the ontology's relevance amidst the fast-paced advancements in AI but also significantly enhances its utility for researchers, developers, and educators by simplifying the integration of new AI concepts and methodologies. The ontology's utility is demonstrated through the annotation of AI methods data in a catalog of AI research publications and the integration into the BioPortal ontology resource, highlighting its potential for cross-disciplinary research. The AIO ontology is open source and is available on GitHub ( https://w3id.org/aio/ ) and BioPortal ( https://bioportal.bioontology.org/ontologies/AIO ).

Joachimiak, Marcin P. [Biosystems Data Science Dep↗

Genome-wide profiling of histone (H3) lysine 4 (K4) tri-methylation (me3) under drought, heat, and combined stresses in switchgrass

Background: Switchgrass (Panicum virgatum L.) is a warm-season perennial (C4) grass identified as an important biofuel crop in the United States. It is well adapted to the marginal environment where heat and moisture stresses predominantly affect crop growth. However, the underlying molecular mechanisms associated with heat and drought stress tolerance still need to be fully understood in switchgrass. The methylation of H3K4 is often associated with transcriptional activation of genes, including stress-responsive. Therefore, this study aimed to analyze genome-wide histone H3K4-tri-methylation in switchgrass under heat, drought, and combined stress. Results: In total, ~ 1.3 million H3K4me3 peaks were identified in this study using SICER. Among them, 7,342; 6,510; and 8,536 peaks responded under drought (DT), drought and heat (DTHT), and heat (HT) stresses, respectively. Most DT and DTHT peaks spanned 0 to + 2000 bases from the transcription start site [TSS]. By comparing differentially marked peaks with RNA-Seq data, we identified peaks associated with genes: 155 DT-responsive peaks with 118 DT-responsive genes, 121 DTHT-responsive peaks with 110 DTHT-responsive genes, and 175 HT-responsive peaks with 136 HT-responsive genes. We have identified various transcription factors involved in DT, DTHT, and HT stresses. Gene Ontology analysis using the AgriGO revealed that most genes belonged to biological processes. Most annotated peaks belonged to metabolite interconversion, RNA metabolism, transporter, protein modifying, defense/immunity, membrane traffic protein, transmembrane signal receptor, and transcriptional regulator protein families. Further, we identified significant peaks associated with TFs, hormones, signaling, fatty acid and carbohydrate metabolism, and secondary metabolites. qRT-PCR analysis revealed the relative expressions of six abiotic stress-responsive genes (transketolase, chromatin remodeling factor-CDH3, fatty-acid desaturase A, transmembrane protein 14C, beta-amylase 1, and integrase-type DNA binding protein genes) that were significantly (P < 0.05) marked during drought, heat, and combined stresses by comparing stress-induced against un-stressed and input controls. Conclusion: Our study provides a comprehensive and reproducible epigenomic analysis of drought, heat, and combined stress responses in switchgrass. Significant enrichment of H3K4me3 peaks downstream of the TSS of protein-coding genes was observed. In addition, the cost-effective experimental design, modified ChIP-Seq approach, and analyses presented here can serve as a prototype for other non-model plant species for conducting stress studies.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Omics integration can be used to rescue metabolic information for some of the dark region of the Pseudomonas putida proteome

In every omics experiment, genes or their products are identified for which even state of the art tools are unable to assign a function. In the biotechnology chassis organism Pseudomonas putida, these proteins of unknown function make up 14% of the proteome. This missing information can bias analyses since these proteins can carry out functions which impact the engineering of organisms. As a consequence of predicting protein function across all organisms, function prediction tools generally fail to use all of the types of data available for any specific organism, including protein and transcript expression information. Additionally, the release of Alphafold predictions for all Uniprot proteins provides a novel opportunity for leveraging structural information. We constructed a bespoke machine learning model to predict the function of recalcitrant proteins of unknown function in Pseudomonas putida based on these sources of data, which annotated 1079 terms to 213 proteins. Among the predicted functions supplied by the model, we found evidence for a significant overrepresentation of nitrogen metabolism and macromolecule processing proteins. These findings were corroborated by manual analyses of selected proteins which identified, among others, a functionally unannotated operon that likely encodes a branch of the shikimate pathway.

60 APPLIED LIFE SCIENCES↗

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Constitutive Down-Regulation of Liguleless Alleles in Sorghum Drives Increased Productivity and Water Use Efficiency"

Plant architecture influences the microenvironment throughout the canopy layer. Plants with a more erect leaf architecture allow for an increase in planting densities and allow more light to reach lower canopy leaves. This is predicted to increase crop carbon assimilation. Frictional resistance to wind reduces air movement in the lower canopy, resulting in higher humidity. By increasing the proportion of canopy photosynthesis in the more humid lower canopy, gains in the efficiency of water use might be expected, although this may be slightly offset by the more open erectophile form canopy. An anatomical feature in members of the Poaceae family that impacts leaf angle is the articulated junction of the sheath and blade, which also bares the ligule and auricles. Mutants, which lack ligules and auricles, show no articulation at this junction, resulting in leaves that are near vertical. In maize, these phenotypes termed liguleless result from null mutations of genes: ZmLG1 (Zm00001eb67740) and ZmLG2 (Zm00001eb147220). In sorghum, SbiRTx430.06G264300 (SbLG1) and SbiRTx430.03G392300 (SbLG2) are annotated as the respective maize homologues. A hair-pin element designed to down-regulate both SbLG1 and SbLG2 was introduced into the grain sorghum genotype RTx430. Derived transgenic events harbouring the hair-pin failed to develop ligules and displayed reduced leaf angles to the vertical, but less vertical than in null mutations. Under field settings, plots sown with these sorghum events having an erect architecture phenotype displayed an increase in photosynthesis in lower canopy levels, which led to increases in above-ground biomass and seed yield, without an increase in water use.

Genome Engineering↗