Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Experimental validation of a collision-radiation dataset for molecular hydrogen in plasmas

Quantitative spectroscopy of molecular hydrogen has generated substantial demand, leading to the accumulation of diverse elementary process data encompassing radiative transitions, electron-impact transitions, predissociations, and quenching. However, their rates currently available are still sparse, and there are inconsistencies among those proposed by different authors. In this study, we demonstrate an experimental validation of such a molecular dataset by composing a collisional-radiative model (CRM) for molecular hydrogen and comparing experimentally obtained vibronic populations across multiple levels. From the population kinetics of molecular hydrogen, the importance of each elementary process in various parameter space is studied. In low-density plasmas (electron density ne≲1017 m−3) the excitation rates from the ground states and radiative decay rates, both of which have been reported previously, determine the excited state population. The inconsistency in the excitation rates affects the population distribution the most significantly in this parameter space. However, in higher density plasmas (ne≳1018 m−3), the excitation rates from excited states become important, which have never been reported in the literature, and may need to be approximated in some way. In order to validate these molecular datasets and approximated rates, we carried out experimental observations for two different hydrogen plasmas; a low-density radio frequency heated plasma (ne≈1016 m−3) and the Large Helical Device (LHD) divertor plasma (ne≳1018 m−3). The visible emission lines from EF1Σg+, HH¯1Σg+, D1Πu±, GK1Σg+, I1Πg±, J1Δg±, h3Σg+, e3Σu+, d3Πu±,g3Σg+, i3Πg±, and j3Δg± states were observed simultaneously and their population distributions were obtained from their intensities. We compared the observed population distributions with the CRM prediction, in particular the CRM with the rates compiled by Janev et al., Miles et al., and those calculated with the molecular convergent close-coupling (MCCC) method. The MCCC prediction gives the best agreement with the experiment, particularly for the emission from the low-density plasma. However, the population distribution in the LHD divertor shows a worse agreement with the CRM than those from low-density plasma, indicating the necessity of the precise excitation rates from excited states. We also found that the rates for the electron attachment is inconsistent with experimental results. This requires further investigation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Dark energy survey: Modeling strategy for multiprobe cluster cosmology and validation for the full six-year dataset

Here, we introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3×2 pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (i) identification of modeling and parametrization choices and impact studies using simulated analyses with each possible model misspecification and (ii) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3×2 pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy–galaxy lensing and cosmic shear two-point statistics) and the CL+3×2 pt+BAO+SN (combination of CL+3×2 pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

79 ASTRONOMY AND ASTROPHYSICS↗

Tethys Water Demand Data

U.S. water demand varies sharply by sector and region as land use, population, weather patterns, and economic activity co-evolve. High-resolution water demand data is required to capture these dynamics, support integrated energy-water-land modeling, and local-to-regional water scarcity assessments. This dataset contains gridded (1/8 degree), monthly, multi-sector water demand dataset for the contiguous United States (CONUS) covering 1980-2100 across eight future scenarios of human-Earth system change. The dataset covers irrigation, thermoelectric, municipal (public-supply and domestic), livestock, manufacturing, and mining demands, separately for withdrawals and consumption, and includes per-cell renewable vs. non-renewable water source attributions. The dataset is validated against the latest USGS 2010-2020 water-use data for the three largest water demand sectors (Domestic, Electricity, and Irrigation), with correlations ranging from 0.73-0.95 at the HUC6 scale. The two datasets largely agree on an aggregate basis with per-sector bias falling within +/-7%, but they disagree on the spatial allocation of water with individual HUC6 basins having normalized RMSE from 68-171% and median absolute percent difference from 37-86%. This dataset advances prior global products by combining state-resolved sectoral demands from GCAM-USA, future power-plant siting from the CERF model, and scenario-consistent high-resolution climate and population forcing data across the eight scenarios.

GCAM-USA↗

Thermal Performance of Spandrel Assemblies in Glazed Wall Systems: Laboratory Test Design – Challenges and Test Results

Accurate thermal performance calculation procedures for opaque spandrel areas in curtain wall and window wall systems are essential for rating systems when comparing spandrel systems. However, there is a lack of consensus in thermal modeling needed for accurately characterizing heat transfer through spandrel assemblies due to the complex arrangement of materials and structural components. Several studies indicate that conventional 2D thermal simulations may overestimate R-values by 30% compared to physical testing and 3D simulations. Detailed simulations and well-curated laboratory test data are necessary to build confidence in simulation models, which will later be used to develop correlations to improve widely used conventional 2D thermal simulations. This study aims to experimentally test heat transfer through various spandrel assemblies to validate 3D simulation models. Also, the challenges of conducting a thorough testing design along with the solutions would be documented. The team developed a design for testing spandrel assemblies, making appropriate modifications to the existing heat, air, and moisture (HAM) chamber to accommodate the testing needs. Two moveable baffles were designed and fabricated to guide airflow direction parallel to the test article surface. The data acquisition capabilities in the chamber were upgraded to add more than two hundred sensors to the climate and indoor side of the chamber. The goal is to provide a quality dataset for validating complex 3D modeling simulations, which will be used to develop improved thermal simulation techniques that more accurately represent the thermal behavior of spandrel assemblies and their integration within the building envelope. This paper will summarize the results for the boundary conditions of the testing and the temperature variation across different locations of the spandrel assemblies.

Kunwar, Niraj [ORNL] (ORCID:0000000263457652)↗

Location-Specific Microstructure Characterization Within AM Bench 2022 Nickel Alloy 718 3D Builds

Abstract The Additive Manufacturing Benchmark Test Series (AM Bench) is a broad effort to produce rigorous measurement datasets for validating AM computer simulations across the range of processing, structure, and properties, for many additive manufacturing (AM) build methods and material classes. Here, the microstructures of nickel alloy 718 AM Bench 2022 test artifacts produced using laser-based powder bed fusion (PBF-LB), in both as-built and fully heat-treated conditions, are examined. Cross sections are primarily characterized using large area scanning electron microscopy (SEM) electron backscatter diffraction (EBSD) and example analyses of the crystallographic textures are described. These data are part of a large set of in situ and ex situ measurements from both three-dimensional builds and laser tracks on bare plates. All the measurement data are available online with download links at www.nist.gov/ambench .

Levine, L. E. (ORCID:0000000334484229)↗

Nuclear quantum effects of metal surface-mediated C–H activation

The nuclear quantum effects of surface-mediated C–H activation of surface CH 3 are considered for the pristine Pt(111) and Au(111) surfaces at 300 K. The kinetic barriers without nuclear quantum effects are calculated using both static density functional theory calculations and ab initio molecular dynamics. Static calculations are performed using the harmonic approximation while the free energy pathway is calculated using enhanced sampling molecular dynamics. Machine learning potentials are trained using generated datasets and validated against the ab initio molecular dynamics generated free energy pathways. The machine learning potentials are used to perform centroid molecular dynamics to consider the nuclear quantum effects of C–H activation. Nuclear quantum effects are found to have a very significant effect on the free energy pathway, with reduced importance at higher temperatures and in the CD 3 case.

Bunting, Rhys J. [Lawrence Livermore National Labo↗

Bayesian Gaussian process inference for neutron spin echo measurement

Neutron spin echo (NSE) spectroscopy provides unique access to microscopic dynamics, but its application is often constrained by low neutron flux, long acquisition times, and significant noise. Here, we present a Bayesian inference approach based on Gaussian process regression (GPR) to reconstruct high-quality spin echo signals from sparse and noisy data by exploiting correlations in reciprocal space. Benchmarks on synthetic datasets and validation with experimental NSE measurements of dendrimers show that GPR suppresses noise, interpolates missing intensity values, and accommodates irregular observations. The method improves accuracy, shortens acquisition times, and enables high-throughput and real-time studies. Beyond NSE, the framework is broadly applicable to other low signal-to-noise ratio scattering techniques, thereby extending the scope of neutron spectroscopy.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

First DIII-D-West hybrid scenario similarity experiments for iter-relevant long-pulse operation

For the first time, similarity experiments between DIII-D and WEST were performed in the ITER "hybrid-like" regime during dedicated campaigns in April and May 2025. The matched parameters include elongation, triangularity, ion ∇B drift direction toward the X-point, qprofile, and core normalized physics quantities in terms of normalized pressure, normalized gyroradius, electron collisionality, ratio of ion to electron temperature, T i /T e . Core transport physics is explored with different aspect ratio (R/a) values (typically 3 at DIII-D and 5 on WEST). DIII-D explored high-beta conditions (electromagnetic effect) with low torque injection (~0 ± 0.5 N•m) using high heating power (up to 6 MW NBI and 2 MW ECRH powers), while scanning the heating mix (ion vs electron), beta, T i /T e , core radiation via controlled tungsten injection using the Laser Blow-Off system. WEST extended operation toward long-duration pulses using its actively cooled tungsten divertor, achieving dominated electron heating regimes with reduced tungsten contamination. Boron impurity injection were scanned on WEST to control edge conditions and core performance. It is found that core confinement improves-manifested by higher electron temperature, total energy content, neutron rate, and ion temperatureunder conditions of low separatrix density, consistent with previous observations [Bourdelle et al., Nucl. Fusion 63 (2023) 056021]. Conditions for Hmode access and for ion heating in electron-dominated regimes in both WEST and DIII-D will be discussed and compared. The ratio of the thermal energy confinement time (τ E ) to the volume-averaged electron-ion collisional heat exchange time (τ e-i ) is a key parameter to enhance ion heating and potentially facilitate H-mode access in electron-heated regimes. These first-of-a-kind coordinated DIII-D and WEST experiments provide a unique multi-machine dataset to validate predictive models and to optimize ITER hybrid-scenario performance under diverse core and edge conditions.

DIII-D↗

An international benchmark for wind plant wakes from the American WAKE ExperimeNt (AWAKEN)

This article introduces the first benchmark study within the International Energy Agency Wind Task 57 framework, focusing on wind plant wakes. Leveraging data from the American WAKE ExperimeNt (AWAKEN), the benchmark aims to assess the accuracy of simulation tools in modeling wind plant wakes and their impact on the downstream flow under diverse inflow conditions. The AWAKEN field campaign, conducted in Oklahoma from 2022 to 2024, provides unprecedented observations of wind plant-atmosphere interactions, thus offering a large dataset to validate numerical models of different complexity. The benchmark will include three phases—code calibration, blind comparison, and iteration—allowing participants to refine their numerical models based on the feedback from the benchmark team. This article describes the benchmark case study selected from observations providing details on atmospheric conditions, wake evidence, and wind turbine operation. The benchmark’s structure and timeline, along with the expected publication of results, are discussed as well. This collaborative effort aims to enhance the accuracy of wind plant wake simulations, thus contributing to the improvement of wind energy production estimates.

17 WIND ENERGY↗

OSW Consortium 2 - Validated National Offshore Wind Resource Dataset with Uncertainty Quantification (CRADA Report)

This research has led to the development of the 2023 National Offshore Wind data set (NOW-23), which offers the latest wind resource information for offshore regions in the United States. NOW-23 supersedes, for its offshore component, the Wind Integration National Dataset (WIND) Toolkit, which was published a decade ago and is currently a primary resource for wind resource assessments and grid integration studies in the contiguous United States. By incorporating advancements in the Weather Research and Forecasting (WRF) model, NOW-23 delivers an updated and cutting-edge product to stakeholders. As part of this project, we also developed a summary of the uncertainty quantification in NOW-23, along with NOW-WAKES, a 1-year post-construction data set that quantifies expected offshore wake effects in the US Mid-Atlantic lease areas. Stakeholders can access the NOW-23 data set at https://doi.org/10.25984/1821404.

17 WIND ENERGY↗

PLUSWIND Derived Data

This dataset consists of annual CSV files containing multiple sources of modeled, hourly wind speeds and generation. For complete information about this dataset, including validation of modeled generation versus recorded generation, please see the Scientific Data article: Millstein, D., Jeong, S., Ancell, A., & Wiser, R. (2023). A database of hourly wind speed and modeled generation for US wind plants based on three meteorological models. Scientific Data, 10(1), 883. https://doi.org/10.1038/s41597-023-02804-w

17 WIND ENERGY↗

Seasonal Precipitation Classification during Surface Atmosphere Integrated Field Laboratory Campaign

The Surface Atmosphere Integrated Field Laboratory (SAIL) campaign, conducted from September 2021 to June 2023 in Crested Butte, Colorado, aimed to characterize precipitation processes in the Upper Colorado River Basin (UCRB). This increased observations of snowfall accumulation in this hydrologically significant watershed would be useful for quantitative precipitation estimates (QPE). Therefore, the Surface Quantitative Precipitation Estimate (SQUIRE) product was developed using the ARM-supported Colorado State University (CSU) X-band Precipitation Radar. Although SQUIRE will be only released for snowfall, by categorizing precipitation types, users can effectively utilize relevant datasets under diverse meteorological conditions. Moreover, the dataset facilitates validation of the QPE product and the analysis of seasonal variations in precipitation types at the surface. Hydrometeors classes are organized based on their phase and physical characteristics mapping the CSU (both winter Summer) and Py-ART classifications into four groups. 1. Liquid Precipitation: includs drizzle, rain, and large raindrops. 2. Frozen Snow and Ice : Pure Snow, combining ice crystals, aggregates, and vertically oriented ice structures. 3. Dense and Large frozen hydrometeors: including low- and high-density graupel and dry hail. 4.Melting: Wet Snow and Melting Hail, hydrometeors exhibiting both liquid and frozen characteristics.

54 ENVIRONMENTAL SCIENCES↗

Dataset for ASME VVUQ Symposium Workshop on Regression of Validation Data to an Application Point

This dataset consists of a collection of Excel spreadsheets that contain output from analysis specified in the workshop. The analysis involves ASME V&V 20-style validation as well as the application of a supplement methodology for regression of validation comparison error and validation uncertainty to application points where experimental data does not exist for comparison. The simulation results and experimental data are provided by the workshop organizers and a NASA report, respectively.

Kirsch, Jared Roelof [Sandia National Laboratories↗

Multi-Artifact Analysis of Self-Admitted Technical Debt in Scientific Software

Context: Self-admitted technical debt (SATD) occurs when developers acknowledge shortcuts in code. In scientific software (SSW), such debt poses unique risks to the validity and reproducibility of results. Objective: This study aims to identify, categorize, and evaluate scientific debt, a specialized form of SATD in SSW, and assess the extent to which traditional SATD categories capture these domain-specific issues. Method: We conduct a multi-artifact analysis across code comments, commit messages, pull requests, and issue trackers from 23 open-source SSW projects. We construct and validate a curated dataset of scientific debt, develop a multi-source SATD classifier to guide SATD management, and conduct a practitioner validation to assess the practical relevance of scientific debt. Results: Our classifier performs strongly across 900,358 artifacts from 23 SSW projects. SATD is most prevalent in pull requests and issue trackers, underscoring the value of multi-artifact analysis. Models trained on traditional SATD often miss scientific debt, emphasizing the need for its explicit detection in SSW. Practitioner validation confirmed that scientific debt is both recognizable and useful in practice. Conclusions: Scientific debt represents a unique form of SATD in SSW that that is not adequately captured by traditional categories and requires specialized identification and management. Our dataset, classification analysis, and practitioner validation results provide the first formal multi-artifact perspective on scientific debt, highlighting the need for tailored SATD detection approaches in SSW.

Melin, Eric [Boise State University]↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Open Source Synergy: Developing and Validating PMU Data Analysis Techniques Using Open Source Tools and Datasets

This paper presents an exploration into the development and validation of data analysis approaches for Phasor Measurement Units (PMUs) using open-source datasets and tools. Various methods for event detection, event classification, frequency response, and oscillation analysis were tested. We leverage the capabilities of Archive Walker (AW), the Frequency Response Analysis Tool (FRAT), and the Oscillation Baselining and Analysis Tool (OBAT), all open-source tools, for efficient processing and analysis of synchrophasor data. The open-source Transmission Signature Library (TSL) dataset was employed as a dataset for a comprehensive evaluation to assess the performance and reliability of the proposed methods.

PMU, event analysis, oscillation, Frequency Respon↗

Verified, Archived, Library of Inputs and Data (VALID) Supporting Files

This dataset contains input, output, and sensitivity data files for computational simulations with the SCALE code system as part of the Verified, Archived Library of Inputs and Data (VALID). The simulations cover critical benchmark experiments from the International Criticality Safety Benchmark Evaluation Project. The files are to be housed in a public directory for distribution. The information contained in the files have been approved for release by the Organisation for Economic Co-operation and Development Nuclear Energy Agency (NEA). Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases.

keff↗

Characterizing Seasonal Variation of the Atmospheric Mixing Layer Height Using Machine Learning Approaches

As machine learning becomes more integrated into atmospheric science, XGBoost has gained popularity for its ability to assess the relative contributions of influencing factors in the atmospheric boundary layer height. To examine how these factors vary across seasons, a seasonal analysis is necessary. However, dividing data by season reduces the sample size, which can affect result reliability and complicate factor comparisons. To address these challenges, this study replaces default parameters with grid search optimization and incorporates cross-validation to mitigate dataset limitations. Using XGBoost with four years of data from the atmospheric radiation measurement (ARM) (Southern Great Plains (SGP) C1 site, cross-validation stabilizes correlation coefficient fluctuations from 0.3 to within 0.1. With optimized parameters, the R value can reach 0.81. Analysis of the C1 site reveals that the relative importance of different factors changes across seasons. Lower tropospheric stability (LTS, ~0.53) is the dominant factor at C1 throughout the year. However, during DJF, latent heat flux (LHF, 0.44) surpasses LTS (0.22). In SON, LTS (0.58) becomes more influential than LHF (0.18). Further comparisons among the four long-term SGP sites (C1, E32, E37, and E39) show seasonal variations in relative importance. Notably, during JJA, the differences in the relative importance of the three factors across all sites are lower than in other seasons. This suggests that boundary layer development in the summer is not dominated by a single factor, reflecting a more intricate process likely influenced by seasonal conditions such as enhanced convective activity, higher temperatures, and humidity, which collectively contribute to a balanced distribution of parameter impacts. Furthermore, the relative importance of LTS gradually increases from morning to noon, indicating that LTS becomes more significant as the boundary layer approaches its maximum height. Consequently, the LTS in the early morning in autumn exhibits greater relative importance compared to other seasons. This reflects a faster development of the mixing layer height (MLH) in autumn, suggesting that it is easier to retrieve the MLH from the previous day during this period. The findings enhance understanding of boundary layer evolution and contribute to improved boundary layer parameterization.

54 ENVIRONMENTAL SCIENCES↗