Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Event Log / Raw Data

The WFIP3 event log is a curated record spanning 578 days of meteorological phenomena and field observations that complements the campaign’s high-frequency measurements. The log combines manually documented daily weather discussions with automatically derived indicators of key atmospheric processes, providing standardized, publicly available context to support model evaluation, forecast verification, and case-study selection for offshore boundary-layer research.

17 WIND ENERGY↗

Pu(IV) quantification via visible–near-infrared absorption spectroscopy: tackling interferences using D-optimal design and partial least squares

Here, this study presents a novel analytical approach for quantifying Pu(IV) in glove box environments using fiber-optic-based visible–near-infrared absorption spectroscopy in combination with partial least squares regression (PLSR) and design of experiments. The method addresses significant challenges posed by overlapping spectral features arising from Nd(III), which is a common fission product impurity, and the speciation variability of Pu(IV) nitrato complexes in HNO 3 concentrations ranging from 2.5 to 11 M. A curated training set consisting of data from 20 samples was developed via D-optimal design to enable robust PLSR model calibration for Pu(IV) using the near-infrared band near 1050 nm. The training set was acquired from samples in cuvettes with a 1-cm path length and was used to build the PLSR model. The robustness of the model was validated with data collected using a dip probe with a 1-cm path length and varying Pu(IV) concentrations. The strong performance of the model indicates good model transfer from cuvette to dip probe and highlights the potential for in situ measurements and online monitoring of reactions in a crystallization reactor vessel. The results demonstrate that this combined spectroscopic and chemometric approach can accurately and simultaneously quantify Pu(IV) and HNO 3 , thereby offering a promising tool for real-time monitoring in process environments.

Actinide↗

Gold-Standard Chemical Database 137 (GSCDB137): A Diverse Set of Accurate Energy Differences for Assessing and Developing Density Functionals

We present GSCDB137, a rigorously curated benchmark library of 137 data sets (8377 entries) covering main-group and transition-metal reaction energies and barrier heights, (intra- and intermolecular) noncovalent interactions, dipole moments, polarizabilities, electric-field response energies, and vibrational frequencies. Legacy data from GMTKN55 and MGCDB84 have been updated to today's best reference values; redundant or low-quality points were removed, and many new, property-focused sets were added. Testing 29 popular density functional approximations (DFAs) confirms the expected Jacob's-ladder hierarchy overall but also reveals notable exceptions: functional performance for frequencies and electric-field properties correlates poorly with that for other ground-state energetics. ωB97M-V and ωB97X-V are the most balanced hybrid meta-GGA and hybrid GGA, respectively; B97M-V and revPBE-D4 lead the meta-GGA and GGA classes. Double hybrids lower mean errors by about 30% versus their hybrid analogues but demand careful frozen-core, basis set, and spin contamination treatment. GSCDB137 offers a comprehensive, openly documented platform for rigorous validation of DFA and universal machine learning potentials, and training of the next generation of exchange-correlation functionals.

Liang, Jiashu [University of California, Berkeley,↗

Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge

Understanding the interactions and regulatory relationships among biomolecules is essential for deciphering complex biological systems and elucidating the mechanisms behind diverse biological functions. Traditionally, the collection of such molecular interaction data has relied on expert curation, a process that is both time-consuming and labor-intensive. To address these limitations, this study explores the use of large language models (LLMs) to automate the genome-scale extraction of molecular interaction knowledge. Here, we evaluate the performance of various LLMs on key biological tasks, including the identification of protein-protein interactions, detection of genes associated with pathways influenced by low-dose radiation, and inference of gene regulatory relationships. Our findings demonstrate that larger LLMs tend to perform better, particularly in extracting intricate gene and protein interactions. Despite their strengths, these models face challenges in recognizing functionally diverse gene groups and highly correlated regulatory relationships. Through a comprehensive analysis using established molecular interaction and pathway databases, we show that LLMs possess the potential to identify relevant biomolecules and predict their interactions, offering valuable insights and marking a significant step toward AI-driven biological knowledge discovery.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

A Review of Resilience and Long-Term Planning in Power and Water Systems in the United States

There is recognition among power and water utilities that the frequency and magnitude of high consequence and low probability events could increase as a result of climate change. The interconnected nature of energy-water systems raises the possibility of cascading failures, increasing complexity and risks. Resilience and long-term planning are important ways of weathering the effects of climate change. First, to understand more about resilience, we reviewed existing literature on resilience definitions, metrics, and modeling, focusing on integrated water-power systems. Second, to understand how resilience and planning are being applied in practice, we interviewed utilities and organized, curated, and synthesized the interview data to arrive at several key findings, which are presented here. We found that there is not a consistent definition for resilience, yet it is something that utilities regularly plan for, often with different names and varying methods/measures. However, there is a tangible shift in the industry towards defining and determining measurable resilience metrics. While the exact metrics are a work in progress, utilities are taking steps forward by (1) putting people and culture at the center of resilience, (2) recognizing their own interdependencies, and (3) pursuing better cross-sector collaboration.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Dark Energy Survey Year 6 Results: Photometric Dataset for Cosmology

We describe the photometric dataset assembled from the full 6 yr of observations by the Dark Energy Survey (DES) in support of static-sky cosmology analyses. DES Y6 Gold is a curated dataset derived from DES Data Release 2 (DR2) that incorporates improved measurement, photometric calibration, object classification and value-added information. Y6 Gold comprises nearly 5000 deg$^{2}$ of grizY imaging in the south Galactic cap and includes 669 million objects with a depth of i$_{AB}$ ∼ 23.4 mag at a signal-to-noise ratio ∼ 10 for extended objects and a top-of-the-atmosphere photometric uniformity <2 mmag. Y6 Gold augments DES DR2 with simultaneous fits to multiepoch photometry for more robust galaxy shapes, colors, and photometric redshift estimates. Y6 Gold features improved morphological star–galaxy classification with an efficiency of 98.6% and a contamination of 0.8% for galaxies with 17.5 < i$_{AB}$ < 22.5. Additionally, it includes per-object quality information, and accompanying maps of the footprint coverage, masked regions, imaging depth, survey conditions, and astrophysical foregrounds that are used for cosmology analyses. After quality selections, benchmark samples contain 448 million galaxies and 120 million stars. This publication is complemented by data access and documentation.

79 ASTRONOMY AND ASTROPHYSICS↗

Ten questions on future and extreme weather data for building simulation and analysis in a changing climate

Weather plays a significant role in building operations as it directly influences HVAC loads and in turn the building energy and thermal performance. In a changing climate, future trends and extreme weather events become critical concerns in the global building decarbonization and clean energy transition. This paper aims to address ten key questions concerning extreme and future weather data for building applications, and more importantly to identify research gaps and guide the curation and selection of future and extreme weather data for use in building performance simulation and assessment. The paper intends to inform architects and engineers, operators, owners, policy makers, and other stakeholders on considering the impacts of future and extreme weather data and adopting strategies for selecting and applying this data in various use cases related to building design, operation, and retrofit for energy efficiency, electrification, and climate resilience.

Yan, Da↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

From RNNs to Foundation Models: An Empirical Study on Commercial Building Energy Consumption

Accurate short-term energy consumption forecasting for commercial buildings is crucial for smart grid operations. While smart meters and deep learning models enable forecasting using past data from multiple buildings, data heterogeneity from diverse buildings can reduce model performance. The impact of increasing dataset heterogeneity in time series forecasting, while keeping size and model constant, is understudied. We tackle this issue using the ComStock dataset, which provides synthetic energy consumption data for U.S. commercial buildings. Two curated subsets, identical in size and region but differing in building type diversity, are used to assess the performance of various time series forecasting models, including finetuned open-source foundation models (FMs). The results show that dataset heterogeneity and model architecture have a greater impact on post-training forecasting performance than the parameter count. Moreover, despite the higher computational cost, finetuned FMs demonstrate competitive performance compared to base models trained from scratch.

commercial buildings↗

SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

The rapid advancement of generative models has made the detection of AI-generated images a critical challenge for both research and society. Recent works have shown that most state-of-the-art fake image detection methods overfit to their training data and catastrophically fail when evaluated on curated hard test sets with strong distribution shifts. In this work, we argue that it is more principled to learn a tight decision boundary around the real image distribution and treat the fake category as a sink class. To this end, we propose SimLBR, a simple and efficient framework for fake image detection with Latent Blending Regularization (LBR). Our method significantly improves cross-generator generalization, achieving up to +24.85% accuracy and +69.62% recall on the challenging Chameleon benchmark. SimLBR is also highly efficient, training orders of magnitude faster than existing approaches. Furthermore, we emphasize the need for reliability-oriented evaluation in fake image detection, introducing risk-adjusted metrics and worst-case estimates to better assess model robustness. All the code and models are availabe at: https://github.com/mvrl/SimLBR

Dhakal, Aayush [Washington University, St. Louis]↗

Antarctic ice sheet model comparison with uncurated geological constraints shows that higher spatial resolution improves deglacial reconstructions

Accurately reconstructing past changes to the shape and volume of the Antarctic ice sheet relies on the use of physically based and thus internally consistent ice sheet modeling, benchmarked against spatially limited geologic data. The challenge in model benchmarking against geologic data is diagnosing whether model-data misfits are the result of an inadequate model, inherently noisy or biased geologic data, and/or incorrect association between modeled quantities and geologic observations. In this work we address this challenge by (i) the development and use of a new model-data evaluation framework applied to an uncurated data set of geologic constraints, and (ii) nested high-spatial-resolution modeling designed to test the hypothesis that model resolution is an important limitation in matching geologic data. While previous approaches to model benchmarking employed highly curated datasets, our approach applies an automated screening and quality control algorithm to an uncurated public dataset of geochronological observations (specifically, cosmogenic-nuclide exposure-age measurements from glacial deposits in ice-free areas). This optimizes data utilization by including more geological constraints, reduces potential interpretive bias, and allows unsupervised assimilation of new data as they are collected. We also incorporate a nested model framework in which high-resolution domains are downscaled from a continent-wide ice sheet model. We highlight the application of this framework by applying these methods to a small ensemble of deglacial ice-sheet model simulations, and demonstrate that the nested approach improves the ability of model simulations to match exposure age data collected from areas of complex topography and ice flow. We develop a range of diagnostic model-data comparison metrics to provide more insight into model performance than possible from a single-valued misfit statistic, showing that different metrics capture different aspects of ice sheet deflation.

Geosciences↗

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗