Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

Data Fusion for the Development of a Multimodal Freight Transload Facilities Dataset in the U.S.

To withstand the growing demand of commodity volume and its strain on the transportation infrastructure, it is necessary to identify the flow of commodities by route and mode. However, a national multimodal freight routing model does not exist for the U.S. The development of such model requires multiple building blocks, such as virtual representations of roadway, railway, and waterway networks, transload facilities (TFs), and access/egress links. Most of these blocks have a robust database in the U.S., except for the TFs. Here, this paper presents the fusion of dispersed and heterogeneous representations of multimodal TFs into a single, comprehensive, geospatial freight TF dataset. The TF dataset is derived from several sources, including the U.S. Army Corps of Engineers Master Docks Plus, the National Transportation Atlas Database, the Intermodal Association of North America, industry publications, and other public information. First, individual datasets were queried and reconciled. A geocoding/reverse geocoding process was applied to get the best street address and latitude/longitude location for each terminal. Then, duplicate terminals were identified by a fuzzy match algorithm based on terminal name and location, and removed. Validation was performed by visual inspection of random facilities. The main contributions of this work are: a publicly available version of the TF dataset, including facility location and multimodal transfer capability of 9,003 facilities, and an enterprise-version with the same facilities but including commodity handling capabilities. The main purpose of developing the TF dataset is to inform multimodal routing algorithms. The proposed TF dataset allows for credibly modeling the multimodal transfer of commodities within shipment routes.

Commodity Routing↗

The ePIC Simulation Campaign Workflow on the Open Science Grid

The ePIC collaboration is realizing the first experiment of the future Electron-Ion Collider (EIC) at the Brookhaven National Laboratory that will allow for a precision study of the nucleons and the nucleus at the scale of sea quarks and gluons through the study of electron-proton/ion collisions. This paper will discuss the current workflow for running centralized simulation campaigns for ePIC on the Open Science Grid (OSG) infrastructure. This involves monthly releases of ePIC software and container deployments to CVMFS, generation of input datasets in HepMC format according to collaboration-defined policy, using Snakemake in CI/CD for validation and benchmarking, and submitting jobs to the OSG condor scheduler for opportunistic running on available resources. File transfers utilize XrootD, and Rucio is used for data management. The workflow is continuously refined to improve daily throughput (currently 50-100k core hours per day) and minimize job failures. Since May 2023, monthly simulation campaigns employing the workflow have cumulatively used over 20 million core hours on the OSG and produced over 350 TB of simulation data. The campaigns incorporate simulations for the broad science program of the EIC and are actively used for the detector and physics studies in preparation of the Technical Design Report (TDR).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Window Observables for Benchmarking Parton Distribution Functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel “window observables” that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-𝑥, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different window observables that can be defined within a region of 𝑥 where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

lattice QCD↗

ComStock Measure Documentation: Interior Lighting Controls

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy’s ResStock™ and ComStock™ models, this work produces national datasets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ComStock Measure Documentation: Fan Static Pressure Reset for Multizone Variable Air Volume Systems

This report assesses the potential for nationwide adoption of a duct static pressure reset in MZ VAV systems in appropriate applications. Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy’s ResStock™ and ComStock™ models, this work produces national datasets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual sub-hourly energy consumption of the commercial building stock across the United States. The “baseline” model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass-adoption impact on the baseline building stock. “Measures” refers to various “what-if” scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public datasets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario—Fan Static Pressure Reset for Multizone Variable Air Volume (VAV) Systems—and briefly introduces key results. The full public dataset can be accessed on the ComStock data lake or via the Data Viewer at comstock.nrel.gov. The public dataset enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ComStock Measure Documentation: Thermostat Setbacks During Unoccupied Periods

This report assesses the potential for nationwide adoption of thermostat setbacks in appropriate applications. Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy’s ResStock™ and ComStock™ models, this work produces national datasets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual sub-hourly energy consumption of the commercial building stock across the United States. The “baseline” model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass-adoption impact on the baseline building stock. “Measures” refers to various “what-if” scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public datasets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario— Thermostat Setbacks During Unoccupied Periods—and briefly introduces key results. The full public dataset can be accessed on the ComStock data lake or via the Data Viewer at comstock.nrel.gov. The public dataset enables users to create custom aggregations of results for their use cases (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Window observables for benchmarking parton distribution functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel ``window observables'' that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-x, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different ``window observables'' that can be defined within a region of x where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Developing machine learning for heterogeneous catalysis with experimental and computational data

Machine learning techniques have emerged as a useful tool for identifying complex patterns and correlations in large datasets, such as associating catalyst performance to its physicochemical properties. In the heterogeneous catalysis communities, machine learning models have mostly been developed using high-throughput quantum chemistry calculations, with only a few case studies resulting in experimentally validated catalyst improvements. This limited success may be due to the use of simplified catalyst structures in computational studies and the lack of comprehensive experimental datasets. In this Review, we bring together studies integrating high-throughput approaches and machine learning for the advancement of solid heterogeneous catalysis, leveraging both experimental and computational data. We systematically analyze trends in the field, based on the descriptors used as model input and output; the materials, devices, or reactions investigated; the dataset size; and the overall achievements. Furthermore, for models reporting unitless R 2 values, we compare the performances based on these mentioned trends.

Computational chemistry↗

Dark Energy Survey: DESI-independent angular BAO measurement

In this work, we present a measurement of the angular baryon acoustic oscillation (BAO) scale from the completed Dark Energy Survey (DES) dataset excluding the area of overlap with the Dark Energy Spectroscopic Instrument (DESI). We follow the same methodology and validation process as in the DES Y6 BAO analysis. We interpret the impact of this measurement in the context of the statistical preference for 𝑤 0 ⁢𝑤 𝑎 cold dark matter (CDM) over Λ⁢CDM when combined with DES Y5 Type Ia supernovae (SN), Planck CMB, and DESI BAO. Based on our previous work, using the full Y6 DES BAO sample, in combination with SN, CMB and DESI data release 1 (DR1) BAO, added 0.3⁢𝜎 in this preference (from 3.7⁢𝜎 to 4.0⁢𝜎), but this ignored possible correlations between datasets. Using our new DESI-independent DES BAO likelihood instead, we find a smaller increase in the statistical preference for 𝑤 0 ⁢𝑤 𝑎 ⁢CDM, from 3.7⁢𝜎 to 3.8⁢𝜎 when using DESI DR1 BAO, and from 4.0⁢𝜎 to 4.1⁢𝜎 when updating to the more recent DESI data release 2 (DR2) BAO. These significances reduce to 3.1⁢𝜎 when using the new calibrated DES SN-Dovekie. Alongside this work, we publicly release baofit_wtheta, the BAO fitting code for the angular correlation function used in the DES Y6 BAO analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Neural Scaling Laws for Jet Generation

Recently observed empirical scaling laws describe the performance of foundation-type models as three independent key quantities -- dataset size, compute, and model parameters -- are modified. Extracting these scaling laws informs the training of large complex models for which the tuning of hyperparameters in traditional ways is not feasible. This work for the first time explores if scaling laws can also be observed for the task of particle jet generation -- both relevant as a pre-training objective for foundation models and as in-situ simulation by itself. We indeed replicate the key logarithmic scaling law behavior for model-size scaling. Beyond studying the next token prediction validation loss of the generative model, we also study the sliced Wasserstein distance of five physical quantities that are not immediately available to the model during training. Our study shows that this quantity is monotonically related to the next token prediction validation loss, meaning that this loss is indeed a good proxy for the physics performance. For the scaling with dataset size and compute, we observe substantially weaker scaling behavior of both the loss and the sliced Wasserstein distance. We analyze this behavior by introducing the concept of a learnable window, and argue that autoregressive next token prediction on jet constituents exhibits comparatively rapid saturation relative to language-model studies. We discuss possible origins of this behavior, including the stochastic nature of QCD radiation and differences between generative and supervised learning tasks in collider physics.

Amram, Oz [Fermilab]↗

Rapid Detection of Anomalies in Battery Energy Storage System Data

Data analytics is pivotal in assessing the technical characteristics and performance of Battery Energy Storage Systems (BESS), underpinning BESS modeling, optimization, and control. However, raw datasets frequently harbor anomalies from measurement errors and equipment malfunctions, impacting BESS reliability and analysis accuracy To address the challenge, this paper presents a novel methodology for the rapid detection of anomalous charge or discharge cycles within BESS operational data, expediting the cleaning process while ensuring data integrity. We’ve collected diverse and comprehensive real-world BESS operational datasets in collaboration with the Electric Power Research Institute and multiple Washington State utilities. These datasets serve dual roles: enabling comprehensive data exploration and analysis for understanding underlying challenges and method development, while also acting as a vital validation resource, demonstrating practical effectiveness. The proposed method detects anomalies and aids in their resolution, improving system performance characterization precision. It also reveals recurring data anomaly sources, offering insights for data collection and handling enhancement. Practitioners can gain valuable insights from the identified anomalous cycles in the real-world datasets along with the investigative process for root cause analyses and essential data cleaning steps.

Crawford, Aladsair J.↗

Benchmarking Cylindrical Blast Wave Theory Against the OSIRIS-REx Sample Return Capsule Reentry

Weak shock theory based on cylindrical blast waves has been used to interpret meteor infrasound, but it has not been systematically benchmarked against a non-ablating hypersonic source with independently known parameters. The objective of this study is not to propose a new theoretical framework, but to evaluate the operational validity of the existing suite of blast radius formulations against a high-fidelity ground truth dataset. The OSIRIS-REx Sample Return Capsule reentry on 24 September 2023 provides such a benchmark because the capsule geometry, trajectory, and infrasound emission points are constrained from mission data and ray tracing, reducing source-side uncertainty associated with ablation. Using observations from 39 infrasound stations, this benchmarking study evaluates six published blast radius (${R}_{0}$ ) formulations and three weak-shock transition coefficients ( C ) within a stratified atmospheric propagation model to predict signal period and peak overpressure. The benchmarking identifies the Sakurai formulation as the best-performing formulation for non-ablating bodies, with the Jones/Plooster formulation performing comparably when a physically appropriate C is adopted. Sakurai and Jones/Plooster yield linear-period median absolute percentage residuals of 9% and 11%, respectively. The period predictions show only weak sensitivity to C at these propagation distances. The Mach-diameter approximation commonly used in meteor studies overestimates ${R}_{0}$ by more than a factor of 3 in the absence of ablation. Finally, these results establish a performance baseline for applying cylindrical blast wave theory to effectively non-ablating hypersonic bodies and demonstrate that the signal period is a robust observable for constraining ${R}_{0}$.

atmospheric entry↗

CHUWD-H v1.0: a comprehensive historical hourly weather database for U.S. urban energy system modeling

Reliable and continuous meteorological data are crucial for modeling the responses of energy systems and their components to weather and climate conditions, particularly in densely populated urban areas. However, existing long-term datasets often suffer from spatial and temporal gaps and inconsistencies, posing great challenges for detailed urban energy system modeling and cross-city comparison under realistic weather conditions. Here we introduce the Historical Comprehensive Hourly Urban Weather Database (CHUWD-H) v1.0, a 23-year (1998-2020) gap-free and quality-controlled hourly weather dataset covering 550 weather station locations across all urban areas in the contiguous United States. CHUWD-H v1.0 synthesizes hourly weather observations from stations with outputs from a physics-based solar radiation model and a reanalysis dataset through a multi-step gap filling approach. A 10-fold Monte Carlo cross-validation suggests that the accuracy of this gap filling approach surpasses that of conventional gap filling methods. Designed primarily for urban energy system modeling, CHUWD-H v1.0 should also support historical urban meteorological and climate studies, including the validation and evaluation of urban climate modeling.

54 ENVIRONMENTAL SCIENCES↗

Repository of HydroSMADE: Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites

This repository presents HydroSMADE—Hydropower Site-level Monthly Availability Data Ensemble, a new open dataset that provides monthly hydropower availability for 1,593 existing and 124,333 potential sites worldwide over the period 1950–2100. The dataset is generated by using a global hydrologic model (Xanthos) with explicit representation of hydropower operation. Specifically, HydroSMADE distinguishes between storage and diversion sites, applies optimized operating rules, and incorporates site-specific characteristics such as generation capacity, maximum turbine flow, and reservoir storage. Driven by bias-corrected meteorological inputs, the data is provided for 30 alternative future scenarios. The scenarios consist of the full factorial combination of three standard CMIP6 atmospheric forcing pathways (SSP1-2.6, SSP3-7.0, and SSP5-8.5) and ten CMIP6 General Circulation Models (GCMs): GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, EC-Earth3, CanESM5, MIROC6, CNRM-ESM2-1, UKESM1-0-LL, and CNRM-CM6-1. The repository contains a total of 122 files: a text file (readme.txt) containing a brief description of the included data, a CSV file containing site attributes, and the remaining 120 files (in CSV) containing site-level monthly hydropower availability. Example Jupyter Notebooks to explore the HydroSMADE dataset are available on GitHub at https://github.com/kamal0013/HydroSMADE More details on the methods and technical validation of HydroSMADE are available in the following paper by the same authors: Chowdhury, A. K., Abeshu, G. W., Zhao, M., Wild, T. B., Hassan, N., Ying, Z., Kim, G. J., Matthew, B., Jonathan, L., & Li, H.-Y. (Submitted). Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites.

Existing and Potential Sites↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

Daily evapotranspiration changes during heatwaves at 32 NEON sites, 2019-2021

This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.

Day of Heatwave↗

Quantifying Epistemic Uncertainty in Binary Classification via Accuracy Gain

ABSTRACT Recently, a surge of interest has been given to quantifying epistemic uncertainty (EU), the reducible portion of uncertainty due to lack of data. We propose a novel EU estimator in the binary classification setting, as the posterior expected value of the empirical gain in accuracy between the current prediction and the optimal prediction. In order to validate the performance of our EU estimator, we introduce an experimental procedure where we take an existing dataset, remove a set of points, and compare the estimated EU with the observed change in accuracy. Through real and simulated data experiments, we demonstrate the effectiveness of our proposed EU estimator.

97 MATHEMATICS AND COMPUTING↗