Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Underwater Target Detection Software Demonstration on the RivGen Turbine

This repository contains data and processing scripts necessary to train the object detection models utilized in the underwater target detection software demonstration on the RivGen turbine project and to produce performance metrics (precision, recall, mAP50, mAP50-95). - Contents - Data consist of "images" and "labels". Each image has an associated label, both share the same time string in its file name (e.g., 2024_05_25_09_01_57.98.jpg and 2024_05_25_09_01_57.98.txt). Time strings have the format %yyyy_%mm_%dd_%HH_%MM_%SS.%3f. Images and labels were curated from 2021 and 2024 smolt outmigration periods at the project site in Igiugig, AK. Images are monochrome 8-bit images of objects (smolt, debris, and other) passing through the field of view of the deployed cameras during various operational stages of the RivGen turbine. Labels are text files indicating the class and bounding polygon of each object in an image. The provided labels use the "YOLO" label format. - Requirements - Python3.8+ is required to install and run the train and validation script. The README.md provides instruction for installing the requirements from the requirements.py file. - Instructions - The "example_train.py" file ingests the provided data, trains a model, and produces model performance metrics at completion. NOTE: model performance metrics will vary from run to run as a consequence of the random selection of training and validation data.

16 TIDAL AND WAVE POWER↗

Utilizing Ontology Structures To Curate the DOE-NETL Carbon Storage Open Database

The specialized ontology for the Carbon Storage Open Database will enable more rapid assignment of appropriate symbology standards for visualization improvements, optimize topical and spatial tagging within keywords, and improve flexibility for utilization in existing data repositories such as EDX. This effort also aims to establish a foundation for utilization of ontologies for organization of other data related to geologic carbon storage in the future.

Martin, Abigail↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

RefAHL: a curated quorum sensing reference linking diverse LuxI-type signal synthases with their acyl-homoserine lactone products

Some bacteria use acyl-homoserine lactone (AHL) signals in quorum sensing, a type of cell-cell communication. Here, we present “RefAHL,” an updated, curated collection of LuxI-type AHL synthases with their AHL products and associated metadata. RefAHL is publicly available as a community resource to help catalog LuxI-type diversity encoded in (meta) genomic data.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-coupled data-driven design of high-temperature alloys

We present a materials design loop, which streamlines physics-coupled machine learning (ML) surrogate models to discover new alloy chemistries with improved properties. The efficacy is demonstrated by discovering a high-temperature alumina-forming austenitic (AFA) stainless steel with enhanced creep, followed by experimental validation. The ML models have been trained using a well-curated, highly consistent experimental dataset augmented with synthetic microstructural features from a computational thermodynamic approach. We have populated a large number of hypothetical AFA alloys to explore the high-dimensional composition space and have predicted their creep properties by providing the same synthetic input features obtained from the trained ML models. Uncertainties from the ML training were taken as thresholds for truncating predicted results to identify alloys with improved or deteriorated creep. Individual elemental compositions have been determined via probability density distribution analysis from the group of alloys at the top and bottom of the predicted creep values for further virtual and experimental validations. In conclusion, we anticipate that this workflow can be applied to screen desired conditions, such as chemistry and processing parameters, in high-dimensional space through physics-guided data analytics.

Alloy design↗

High Resolution Siting Suitability of Various Power Plant Technologies

Energy sector planning models determine the aggregate need for new generation, but these models are typically at the state or regional scale and are not equipped to address the wide range of location- and technology-specific issues that are increasingly a factor in power plant siting. These animations demonstrate the aggregate siting suitability of various power plant technology configurations, considering technology-specific factors that can prohibit development. The data presented is from the GRIDCERF (Geospatial Raster Input Data for Capacity Expansion Regional Feasibility) data package. GRIDCERF is a harmonized, open-source geospatial product that can be used to evaluate siting suitability for renewable and non-renewable power plants in the conterminous United States. The animations presented here demonstrate a curated selection of the full suite of technology configurations available. GRIDCERF provides the necessary inputs for models that simulate power plant siting for regional capacity expansion planning such as the Capacity Expansion Regional Feasibility (CERF) model.

Mongird, Kendall [Pacific Northwest National Labor↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗

Discovery of hybrid chemical synthesis pathways with DORAnet

Developing efficient tools for discovering novel synthesis pathways is essential to advance chemical production methods that maximize the use of resources and energy. We introduce DORAnet (Designing Optimal Reaction Avenues Network Enumeration Tool), an open-source computational framework that addresses key limitations in current computer-aided synthesis planning (CASP) tools. DORAnet integrates both chemical/chemocatalytic (i.e., non-enzymatic) and enzymatic transformations, enabling the discovery of hybrid synthesis pathways. With 390 expert-curated chemical/chemocatalytic reaction rules and 3606 enzymatic rules derived from MetaCyc, it provides extensive flexibility for synthetic chemists and biotechnologists. The framework features customizable network expansion strategies, advanced filtering, and pathway search, ranking, and visualization tools. Validated against known reaction data, DORAnet successfully identified both established and novel synthesis routes for key industrial chemicals. In a case study involving 51 high-volume targets, DORAnet frequently ranked known commercial pathways among the top three results, demonstrating its practical relevance and ranking accuracy, while also uncovering numerous alternative (hybrid) synthesis pathways that were highly ranked.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

NETL Coal Energy Atlas: A Collection of Coal/Energy Related Maps

The NETL Coal Energy Atlas contains a comprehensive collection of coal and energy-related maps and graphics curated by the National Energy Technology Laboratory (NETL) Systems Analysis group. It serves as a living document providing an overview of the U.S. coal and energy sectors. The volume is structurally organized into six key thematic areas. Ultimately, the atlas functions as a modular baseline for data integration, allowing researchers to drill down into specific regional locations or customize geographic base layers for advanced systems analysis.

bituminous coal↗

Development and transferability of neural-network models for plasma-surface interactions

Plasma-surface interactions are increasingly critical to modern technologies; yet, accurate molecular dynamics simulations remain limited by the capabilities of interatomic potentials. Deep Potentials (DPs) promise to revolutionize the field by providing a systematic method for producing accurate interatomic potentials. The primary challenge of DP development is selecting a dataset, which efficiently spans the set of atomic environments one expects to encounter in the subsequent molecular dynamics simulations. The computational cost of density functional theory calculations, which are the typical basis for DP development, makes it impossible to directly verify the quality of a given DP. To address this challenge, we explore the development of a deep-learned interatomic potential, “DeepREBO,” trained to reproduce the behavior of the REBO2 empirical potential, enabling direct validation of training methodology and transferability. Using an active learning framework, we begin with a minimal dataset and iteratively expand it to train a Deep Potential-Smooth Edition model that faithfully reproduces REBO2 results for 25 eV hydrogen bombardment of diamond (001), a particularly challenging case. We show that small, carefully curated datasets can outperform large, unguided ones, with effective models requiring fewer than 15 000 snapshots. Subsequent transferability tests demonstrate that while DeepREBO generalizes well to diamond (111) surfaces, performance degrades for amorphous carbon or higher-energy impacts, highlighting the need for use-case-specific training data. We also evaluate methods to improve short-range repulsion. This study outlines best practices for training robust deep potentials and underscores the importance of dataset design for predictive plasma simulations.

Ab-initio molecular dynamics↗

Genomic fingerprints of the world’s soil ecosystems

Despite the explosion of soil metagenomic data, we lack a synthesized understanding of patterns in the distribution and functions of soil microorganisms. These patterns are critical to predictions of soil microbiome responses to climate change and resulting feedbacks that regulate greenhouse gas release from soils. To address this gap, we assay 1,512 manually curated soil metagenomes using complementary annotation databases, read-based taxonomy, and machine learning to extract multidimensional genomic fingerprints of global soil microbiomes. Our objective is to uncover novel biogeographical patterns of soil microbiomes across environmental factors and ecological biomes with high molecular resolution. We reveal shifts in the potential for (i) microbial nutrient acquisition across pH gradients; (ii) stress-, transport-, and redox-based processes across changes in soil bulk density; and (iii) greenhouse gas emissions across biomes. We also use an unsupervised approach to reveal a collection of soils with distinct genomic signatures, characterized by coordinated changes in soil organic carbon, nitrogen, and cation exchange capacity and in bulk density and clay content that may ultimately reflect soil environments with high microbial activity. Genomic fingerprints for these soils highlight the importance of resource scavenging, plant-microbe interactions, fungi, and heterotrophic metabolisms. Across all analyses, we observed phylogenetic coherence in soil microbiomes—more closely related microorganisms tended to move congruently in response to soil factors. Collectively, the genomic fingerprints uncovered here present a basis for global patterns in the microbial mechanisms underlying soil biogeochemistry and help beget tractable microbial reaction networks for incorporation into process-based models of soil carbon and nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

The Artificial Intelligence Ontology: LLM-Assisted Construction of AI Concept Hierarchies

The Artificial Intelligence Ontology (AIO) is a systematization of artificial intelligence (AI) concepts, methodologies, and their interrelations. Developed via manual curation, with the additional assistance of large language models (LLMs), AIO aims to address the rapidly evolving landscape of AI by providing a comprehensive framework that encompasses both technical and ethical aspects of AI technologies. The primary audience for AIO includes AI researchers, developers, and educators seeking standardized terminology and concepts within the AI domain. We use the term “branches” for classes, and their subclasses, in our ontology that are subclasses of owl:Thing. AIO contains eight branches: Bias, Layer, Machine Learning Task, Mathematical Function, Model, Network, Preprocessing, and Training Strategy, each designed to support the modular composition of AI methods and facilitate a deeper understanding of deep learning architectures and ethical considerations in AI. AIO uses the Ontology Development Kit (ODK) for its creation and maintenance, with its content being more easily updated through AI-driven curation support. This approach not only ensures the ontology's relevance amidst the fast-paced advancements in AI but also significantly enhances its utility for researchers, developers, and educators by simplifying the integration of new AI concepts and methodologies. The ontology's utility is demonstrated through the annotation of AI methods data in a catalog of AI research publications and the integration into the BioPortal ontology resource, highlighting its potential for cross-disciplinary research. The AIO ontology is open source and is available on GitHub ( https://w3id.org/aio/ ) and BioPortal ( https://bioportal.bioontology.org/ontologies/AIO ).

Joachimiak, Marcin P. [Biosystems Data Science Dep↗

User Guide: A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations

This user guide supports a curated dataset of 71 meteor events recorded between 2006 and 2011 in Southwestern Ontario, Canada. Each event was simultaneously observed by ground-based optical cameras and an infrasound array, providing a rare opportunity to examine meteor trajectories and acoustic signals from the same atmospheric entry events. The dataset includes raw and processed optical data, meteor trajectories, photometric light curves, infrasound waveforms, and atmospheric specifications relevant for acoustic modeling. The archive is structured to support reproducible research in meteor physics, atmospheric acoustics, and shock wave analysis. It is organized following transparent file naming conventions and structured folders to facilitate scientific reuse, comparison, and integration across research domains. The dataset is freely available on Zenodo, doi: 10.5281/zenodo.15868512.

54 ENVIRONMENTAL SCIENCES↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Acoustic Rocket Signatures Collected by Smartphones

Rockets generate complex acoustic signatures that can be detected over a thousand kilometers from their source. While many far-field acoustic rocket signatures have been collected and released to the public, very few signatures collected at distances less than 100 km are available. This work presents a curated and annotated dataset of acoustic signatures of 243 rocket launches collected by a network of smartphones stationed at distances between 10 and 70 km from the launch sites, resulting in 1089 individual recordings. Due to the frequency dependence of atmospheric attenuation and the relatively short propagation distances, higher-frequency features not preserved in most publicly available data are observed. The signals are time-aligned to allow for different segments of the signal (ignition, launch, trajectory, chronology) to be more easily examined and compared. Initial analysis of the features of these rocket launch stages is performed, observed features are compared to those found in the existing literature, and comparisons between signals from launches of different rocket types are made. The dataset is annotated and made available to the public to aid future analysis of the characteristics and source mechanisms of rocket acoustics as well as applications such as rocket detection and classification models.

33 ADVANCED PROPULSION SYSTEMS↗