Search NASASearch

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database

Towards Diverse and Representative Global Pretraining Datasets for Remote Sensing Foundation Models

The design of a pretraining dataset is emerging as a critical component for the generality of foundation models. In the remote sensing realm, large volumes of imagery and benchmark datasets exist that can be leveraged to pretrain foundation models, however using this imagery in absence of a well-crafted sampling strategy is inefficient and has the potential to create biased and less generalizable models. Here, we provide a discussion and vision for the curation and assessment of pretraining datasets for remote sensing geospatial foundation models. We highlight the importance of geographic, temporal, and image acquisition diversity and review possible strategies to enable such diversity at global scale. In addition to these characteristics, support for various spatial-temporal pretext tasks within the dataset is also critical. Ultimately, our primary objective is to place emphasis on and draw attention to the data curation stage of the foundation model development pipeline. By doing so, we think it is possible to reduce biases of geospatial foundation models, as well as enable broader generalization to downstream remote sensing tasks and applications.

Arndt, Jacob

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Agnostic capture of pathogens for the detection and diagnostics of emerging threats

The continued emergence of pathogens, whether novel, re-emerging, or engineered, poses a persistent global biosecurity and public health challenge. Recent outbreaks, including COVID-19, Lassa fever, Marburg virus, mpox, and avian influenza, underscore the urgent need for robust systems that enable rapid surveillance, early diagnosis, and timely countermeasures before widespread human transmission occurs. In this article, we focus on early detection technologies and systematically evaluate current diagnostic and sensing modalities. We highlight sequencing and spectroscopy as two complementary approaches capable of providing broad, agnostic detection and rich biological insight. Our analysis emphasizes that scientific innovation alone is insufficient: effective preparedness also requires improved data curation, integration, and sharing to build AI-ready resources that accelerate future responses. We argue for coordinated advances in both technological capabilities and supporting infrastructure to enable the rapid identification and characterization of emerging pathogens and to fully leverage modern science against evolving infectious threats.

Environmental health

Incorporating Diurnal and Meter-Scale Variations of Ambient CO 2 Concentrations in Development of Direct Air Capture Technologies

To be implemented on climate-relevant scales, direct air capture of CO 2 (DAC) will require large capital-intensive facilities and careful attention to cost minimization. In making decisions among potential sites for DAC facilities, all of the factors that will impact process cost and efficiency should be considered. In this paper we focus on a factor that has previously received little attention in the DAC community, namely variations in atmospheric conditions on hourly time scales and length scales of meters. We present data curated from extensive previous studies of biosphere-atmosphere fluxes with observations of CO 2 concentration, temperature, and relative humidity (RH) with hourly resolution from many sites in North America. These include locations where typical diurnal variations in CO 2 concentration during summer months exceeds 150 ppm. These variations are larger than the seasonal variations that exist between averaged CO 2 concentrations in winter and summer, and they are highly correlated with diurnal variations in temperature and RH. Diurnal variations are dependent on the height above ground at which CO 2 concentrations are measured, with smaller variations existing at heights of 10 m or more than at ground level. We illustrate the potential implications of these short-term variations for the operation and optimization of a DAC process with process-level calculations for a specific adsorption-based process using amine-rich adsorbents.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Advancing the Prediction of MS/MS Spectra Using Machine Learning

Tandem mass spectrometry (MS/MS) is an important tool for the identification of small molecules and metabolites where resultant spectra are most commonly identified by matching them with spectra in MS/MS reference libraries. While popular, this strategy is limited by the contents of existing reference libraries. In response to this limitation, various methods are being developed for the in silico generation of spectra to augment existing libraries. Recently, machine learning and deep learning techniques have been applied to predict spectra with greater speed and accuracy. Here, in this work, we investigate the challenges these algorithms face in achieving fast and accurate predictions on a wide range of small molecules. The challenges are often amplified by the use of generic machine learning benchmarking tactics, which lead to misleading accuracy scores. Curating data sets, only predicting spectra for sufficiently high collision energies, and working more closely with experimental mass spectrometrists are recommended strategies to improve overall prediction accuracy in this nuanced field.

47 OTHER INSTRUMENTATION

A Climatology and Life‐Cycle Characteristics of Atmospheric Fronts and Their Associated Precipitation

Abstract Atmospheric fronts are one of the main sources of mid‐latitude variability. We employ a novel method for identifying and tracking fronts and frontal precipitation. Thermal and dynamical variables are used to identify fronts as areal objects in space, which are tracked in time using the open‐source TempestExtremes software package. Precipitation objects are co‐located to identify frontal precipitation. The method is subjected to validation and sensitivity tests using manually curated data from the National Weather Service. Climatologies of fronts and frontal precipitation are computed from reanalysis and observations; fronts are present upwards of 14% of the time in the storm tracks, and represent the majority (up to 90%) of total and extreme precipitation. Novel aspects of the method are showcased through the lifetime characteristics of fronts across North America. Three sets of warm and cold fronts were discovered, and their duration, distance‐traveled, and translation velocity are examined. Plain Language Summary Mid‐latitude low‐pressure systems and weather fronts are important for our day‐to‐day experience of weather events, particularly in the mid‐latitudes. This work makes use of standardized atmospheric data and creates a method of automatically tracking these important atmospheric features and their precipitation to quantify their relative role in global precipitation. Weather fronts are persistent in the mid‐latitudes and are associated with the majority of precipitation–particularly the most intense precipitation. Trajectories of fronts over North America are categorized to create a set of archetypal fronts that occur in that region. The differences between these types of fronts are characterized. Key Points An automated, efficient, and skillful frontal detection algorithm is developed and validated Fronts contribute a larger fraction of extreme precipitation than all precipitation in mid‐latitude storm tracks Fronts across North America have substantial variation in characteristics depending on their origin location

extratropical cyclone

A Comprehensive Calibration Framework for the Northwest River Forecast Center

We present a comprehensive framework developed by the Northwest River Forecast Center for calibrating hydrologically diverse basins. The framework includes models for snow, soil moisture, routing, channel loss, and consumptive use. Data inputs include a wide range of open-access datasets for meteorology, land use, topography, and land cover. The framework uses conceptual hydrologic models to handle basins with various hydrologic regimes including rain-driven and snowmelt-dominated basins. We also develop a flexible automatic calibration system that can handle numerous unobservable model parameters in a computationally efficient manner. A single-basin automatic calibration run can typically be completed on a modern laptop in under 10 min. We found that model performance metrics for this new approach match the quality of the NWRFC's previous labor-intensive manual calibrations. The model performance also rivals that of a state-of-the-art deep learning model at a fraction of the computational cost. This framework presents a new standard for the quality of calibrations possible with lumped conceptual hydrologic models, combining careful data curation, an objective calibration framework, and expert local knowledge. In addition, we have made software packages available for the entire suite of National Weather Service River Forecast System models, including SAC-SMA, SNOW-17, and Lag-K. These modern interfaces are intended to increase accessibility and facilitate future research.

Forecasting

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 1

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024). Three distinct rounds of FSP experiments were performed by the experimental team, producing replicate samples utilizing across different nominal processing conditions (Condition IDs) listed in Table 1. The starting material on which FSP was applied was commercially available unprocessed stainless-steel type 316L material. Chosen processing conditions were very diverse, and some were intentionally chosen to produce defects. Several samples experienced tool breakage during experimentation, so a full set of three replicates was not produced for every nominal processing condition.

36 MATERIALS SCIENCE

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 2

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024).

36 MATERIALS SCIENCE

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 3

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024).

36 MATERIALS SCIENCE

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 4

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024).

36 MATERIALS SCIENCE

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation

Database of low‐temperature absorption and fluorescence spectra of native photosynthetic tetrapyrrole macrocycles

Low-temperature (77 K) absorption and fluorescence spectra of 12 naturally occurring photosynthetic tetrapyrrole macrocycles have been recorded in a frozen glass (2-methyltetrahydrofuran). The compounds encompass distinct chromophore classes: porphyrin, chlorophyll c 2 ; chlorin, chlorophylls a, b, d, f and bacteriochlorophylls c, d, e, f; and bacteriochlorin, bacteriochlorophylls a, b, g. The spectra are compared with those of the same pigment in liquid solution (predominantly 2-methyltetrahydrofuran) at room temperature (293 K). The measured Stokes shifts at 77 K across the 12 macrocycles range from ~30 to 300 cm −1 . The spectral data in digital form are made available as part of the PhotochemCAD databases. Literature searches have revealed extensive published data for Chl a (often in biological matrices) but at best rather limited data for less common macrocycles. The availability of a systematic collection of curated spectral data collected at low temperature should be useful for a variety of assessments, including reconstruction of absorption spectra of (bacterio)chlorophyll-containing protein complexes, vibrational analysis of absorption and fluorescence spectra, and calculations where knowledge of energy levels is important.

Niedzwiedzki, Dariusz M. [Washington University in

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND