Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

HydroForecast Long-term: Improving hydropower’s resilience to climate change through accurate climate-scale

With hydrologic patterns and water availability across the globe shifting due to climate change, advancements in hydrologic prediction systems can help significantly reduce the uncertainties that utilities and water supply entities have in their decision making. Understanding and estimating hydrology at the climate scale is critical for managing water resources under changing climate scenarios. This project focuses on integrating state-of-the-art neural network modeling with downscaled climate projections to deliver the reliable water supply projections decades into the future to meet an urgent need from hydropower operators and water utilities. In this Phase 1 DOE SBIR proposal, we developed and validated a theory-guided neural network model, HydroForecast Long-term, for climate-scale hydrology and implemented the model within existing HydroForecast infrastructure. HydroForecast Long-term combines the most accurate streamflow modeling system with a flexible and scalable data architecture to generate water supply projections out to the year 2100. This report illustrates that we have achieved our four objectives: 1) create a prototype of HydroForecast Long-term, building the neural network prediction model, 2) build an automated data input pipeline that processes large amounts of data from the latest global temperature and precipitation climate models; 3) benchmark the accuracy of the hydrologic model over the recent two decades over a large set of diverse basins, and 4) create a set of output visuals and summary metrics informed by customer feedback that connect the data to critical decision points. This work empowers water users to make data-informed decisions supporting a resilient, renewable-powered grid and water system. The results advance the Department of Energy’s mission by addressing critical gaps in water supply planning under climate change.

13 HYDRO ENERGY↗

Analyzing the Impact of Future Weather Data on Energy Consumption in Weatherization Assistant

This study supports the mission of the U.S. Department of Energy’s Weatherization Assistance Program (WAP), which aims to increase the energy efficiency of dwellings and reduce their total residential expenditures. Specifically, we examine how projected future climate conditions may affect residential building energy performance by integrating future weather data into the National Energy Audit Tool (NEAT). Since WAP evaluates the cost-effectiveness of retrofit measures over lifespans of up to 30 years, accounting for evolving climate conditions is increasingly important. To reflect future household energy demands, this study replaces historically based Typical Meteorological Year (TMY3) weather inputs with Future Typical Meteorological Year (fTMY) datasets derived from global climate model (GCM) projections. A simulation-based framework was established to enable NEAT analysis under future weather conditions. This workflow involves converting EPW-format weather files into JSON inputs compatible with NEAT and generating degree-hour metrics needed for load calculations. The fTMY dataset used in this study was developed by Oak Ridge National Laboratory through downscaling of six GCMs under different emission scenarios and covers the period from 2020 to 2100. In contrast, the TMY3 dataset is based on historical weather data from 1961 to 1990. Simulations were conducted for benchmark single-family prototype buildings across ASHRAE climate zones 1–7, which cover all regions of the U.S. except the subarctic Zone 8 in northern Alaska, evaluating both heating and cooling loads under TMY3 and fTMY conditions. Four foundation types were tested, while heating systems were standardized, as NEAT does not differentiate thermal energy load by HVAC system type in its load calculations. Results show that fTMY weather input consistently yield lower heating loads and higher cooling loads across most locations, aligning with expected climate warming trends. Notably, colder regions such as zones 6A, 6B, and 7 experience marked reductions in heating load, while warmer and transitional zones, such as 2A (Lufkin, TX) and 3C (San Francisco, CA), have substantial increases in cooling loads. Although this study does not directly assess the performance of retrofit measures under future climate conditions, it provides a critical foundation for doing so. By quantifying shifts in baseline (i.e., pre-retrofit case) energy loads between historical and future weather files, the study highlights the importance of integrating climate-responsive data into audit tools. These findings will inform future efforts to evaluate the long-term effectiveness and cost-effectiveness of weatherization measures under changing climate conditions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Measurement of the 252 Cf ⁢(sf) prompt fission neutron spectrum utilizing 12 C ⁡(𝑛, 𝑛) and 9 Be ⁢(𝑛, 𝑛) neutron scattering reference measurements

The 252 Cf spontaneous fission (sf), prompt fission neutron spectrum (PFNS) is a fundamental quantity for nuclear physics measurements of neutron-emitting reactions. This energy distribution of neutrons emitted from fission has been considered a neutron data standard for decades and has been utilized as a reference for neutron detection efficiency, validation of Monte Carlo simulations, benchmarking of dosimetry standards, and more. A significant portion of the global collection of nuclear data on neutron-induced reactions is correlated with the 252 Cf ⁢(sf) PFNS. Despite the reliance on this quantity by the nuclear physics community, the historical collection of 252 Cf PFNS measurements display systematic disagreements that are not understood or easily explained. These experimental discrepancies could potentially bias the 252 Cf PFNS Standard evaluation. On top of this, these past experiments frequently employed correlated experimental measurement or analysis methods. The artificial intelligence (AI)/machine learning (ML)-informed californium chi-nuclear data experiment (AIACHNE) project was formed to (a) investigate these discrepancies utilizing AI/ML methods to identify outlying regions of literature data, assign these regions to features of the experiment itself, and perform an improved evaluation of the 252 Cf PFNS and (b) perform a new experimental measurement of this quantity designed to improve upon the existing literature database. Here, in this work, we report on the AIACHNE 252 Cf PFNS experiment utilizing a new analysis method uncorrelated with all previous measurements: neutron efficiency determinations based on elastic neutron scattering on 12 C and 9 Be . This new method provides an independent test of the existing literature data and evaluation of the 252 Cf ⁢(sf) PFNS. The method is described with detailed covariance quantification procedures, as well as a direct discussion of the sources of uncertainty described as requirements in the “Templates” series of papers. The 252 Cf ⁢(sf) PFNS reported in this work agrees well with the overall shape of the existing standard PFNS evaluation as well as many literature measurements, thus verifying the current evaluation utilizing new techniques. However, the results suggest that there are deficiencies in the angle-differential 12 C and 9 Be ⁢(𝑛, 𝑛) evaluated nuclear data, which produce unphysical structures in the reported result. While these structures are relatively minor, they become obvious because of the high statistical precision of the data and the expected smooth continuity of the 252 Cf ⁢(sf) PFNS.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Observational Data for Next-Generation Climate Model Evaluation: Requirements, Considerations, and Best Practices

Climate model simulations are an important source of information about our planet’s climate system and also enable informed decision-making under different future scenarios. As a new archive of results from the next generation of climate models is anticipated to become available with the Coupled Model Intercomparison Project phase 7 (CMIP7), the need to develop efficient and robust methods to evaluate models is paramount. Observations are an integral part of model evaluation, providing a means to quantify and understand the degree to which climate models can faithfully reproduce Earth system processes. Such analysis is critical for constraining climate projections, identifying areas of focus for model development, and assisting analysts in deciphering the utility of models for specific applications. Observations of Earth system come from a diversity of sources, span different space–time domains, and are produced by different communities, and each dataset features different data structures and formats, metadata standards, and its own unique uncertainties. Uncertainties in an observational dataset may stem from gaps in temporal and spatial coverage, instrumentation errors, or assumptions in retrieval and processing methods. How then does one ensure that observational data are ready for use and utilized in the most appropriate way for robust, rapid, and routine climate model evaluation? The CMIP7 Model Benchmarking Task Team with input from the broader climate modeling, model evaluation, and observational data communities present a vision and considerations for best practices toward the optimal and appropriate use of observational data to support next-generation climate model evaluation.

Climate models↗

Evaluation of Sandia NCS Benchmark Suite Updates

The Sandia Nuclear Criticality Safety (NCS) program’s benchmark suite was recently updated. This suite is used to ensure that NCS calculations using computer-based neutron transportation codes have an established baseline comparison of calculated versus known experimental results. The Evaluated Nuclear Data File (ENDF) version used for the MCNP models in the suite was changed from ENDF/B-VII.1 to ENDF/B-VIII.0. Additionally, relevant thermal scattering law data libraries (TSLs) were updated. The sensitivity of the calculational bias of each benchmark model to these changes is discussed. Implementation of the ENDF/B-VIII.0 library and updated TSLs results in improvements to bias distribution in the intermediate enriched uranium, plutonium, and mixed uranium–plutonium (IEU, PU, and MIX) fissionable material benchmark categories, but a small bias increase in low- and high-enriched uranium categories (LEU and HEU, respectively). The results also highlight the sensitivity of the benchmarks, with average lethargy of neutrons causing fission energies (EALF) in the intermediate energy range to ENDF/B library changes. The most numerous bias changes were observed in the thermal energy region when transitioning from ENDF/B-VII.1 to ENDF/B-VIII.0. In conclusion, most of the unique bias changes observed in MCNP 6.3.0 between the two nuclear data libraries were in the LEU-COMP-THERM evaluation subset.

ICSBEP↗

Preliminary Comparison of Neutron Detector Sensitivities [Poster]

Nuclear data (ND) is vital to predictive simulations such as those completed with MCNP®. ND sensitivities are used to optimize the design of benchmark experiments. Diverse benchmarks (including neutron noise) are required to further our understanding of ND. Measured and simulated data of a 4.5-kg Pu sphere was used to calculate neutron noise parameters and associated ND Sensitivities

Nuclear Criticality Safety Program (NCSP)↗

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for this task, yet no systematic evaluation exists of how well vision-language models (VLMs) interpret them. We introduce QCalEval, the first VLM benchmark for quantum calibration plots: 243 samples across 87 scenario types from 22 experiment families, spanning superconducting qubits and neutral atoms, evaluated on six question types in both zero-shot and in-context learning settings. The best general-purpose zero-shot model reaches a mean score of 72.3, and many open-weight models degrade under multi-image in-context learning, whereas frontier closed models improve substantially. A supervised fine-tuning ablation at the 9-billion-parameter scale shows that SFT improves zero-shot performance but cannot close the multimodal in-context learning gap. As a reference case study, we release NVIDIA Ising Calibration 1, an open-weight model based on Qwen3.5-35B-A3B that reaches 74.7 zero-shot average score.

Cao, Shuxiang↗

Transfer learning-based soybean LAI estimations by integrating PROSAIL, UAV, and PlanetScope imagery

Accurate Leaf Area Index (LAI) estimations at the soybean plot scale is achievable using high-resolution Unmanned Aerial Vehicle (UAV) imagery and field measurement samples. However, the limited coverage of UAV flights restricts large-scale remote sensing monitoring in expansive soybean fields. This study leverages the broad coverage and 3-m resolution of PlanetScope satellite imagery to extend LAI prediction from UAV to satellite scales through transfer learning, using UAV-scale LAI estimates as a benchmark to validate cross-scale consistency. To address this challenge, this study proposed the LAI-TransNet, a two-stage transfer learning framework designed for precise and scalable soybean LAI prediction across large areas, demonstrating its effectiveness in cross-scale monitoring. In Stage 1, a UAV-scale benchmark is established using PROSAIL-simulated UAV reflectance data (UAV-Sim) and field-measured soybean LAI. Traditional machine learning, deep learning, and transfer learning models are trained on a hybrid UAV-Sim and field-measured dataset (UAV-Sim_Measured), with the transfer learning model CNN-TL, fine-tuned using pre-trained weights derived from UAV-Sim, achieving the highest accuracy (R 2 = 0.81, RMSE = 0.64 m 2 /m 2 , rRMSE = 11.5 %). In Stage 2, LAI-TransNet is developed by fine-tuning the CNN-TL model on PlanetScope simulated data (PS-Sim), preprocessed via cross-domain mapping to align UAV and satellite spectral features. Real PlanetScope imagery is corrected for reflectance consistency with reference to UAV imagery spectral profiles. LAI-TransNet outperforms other deep learning models trained directly on PS-Sim (R 2 = 0.69 vs. 0.60–0.63), ensuring robust cross-scale consistency. In conclusion, by bridging UAV and satellite scales, LAI-TransNet enables large-scale soybean LAI monitoring, enhancing precision agriculture management through improved monitoring with the PlanetScope imagery.

Leaf area index (LAI)↗

Advancing the Prediction of MS/MS Spectra Using Machine Learning

Tandem mass spectrometry (MS/MS) is an important tool for the identification of small molecules and metabolites where resultant spectra are most commonly identified by matching them with spectra in MS/MS reference libraries. While popular, this strategy is limited by the contents of existing reference libraries. In response to this limitation, various methods are being developed for the in silico generation of spectra to augment existing libraries. Recently, machine learning and deep learning techniques have been applied to predict spectra with greater speed and accuracy. Here, in this work, we investigate the challenges these algorithms face in achieving fast and accurate predictions on a wide range of small molecules. The challenges are often amplified by the use of generic machine learning benchmarking tactics, which lead to misleading accuracy scores. Curating data sets, only predicting spectra for sufficiently high collision energies, and working more closely with experimental mass spectrometrists are recommended strategies to improve overall prediction accuracy in this nuanced field.

47 OTHER INSTRUMENTATION↗

Accuracy of Kohn–Sham density functional theory for warm- and hot-dense matter equation of state

We study the accuracy of Kohn–Sham density functional theory (DFT) for warm- and hot-dense matter (WDM and HDM). Specifically, considering a wide range of systems, we perform accurate ab initio molecular dynamics simulations with temperature-independent local/semilocal density functionals to determine the equations of state at compression ratios of 3x–7x and temperatures near 1 MK. We find very good agreement with path integral Monte Carlo benchmarks, while having significantly smaller error bars and smoother data, demonstrating the accuracy of DFT for the study of WDM and HDM at such conditions. In addition, using a Δ-machine learned force field scheme, we confirm that the DFT results are insensitive to the choice of exchange-correlation functional, whether local, semilocal, or nonlocal.

Suryanarayana, Phanish (ORCID:0000000151720049)↗

Demonstration of TOFFEE: A Response Uncertainty Quantification Tool

A key characteristic in neutron transport is nuclear data. Cross-section uncertainty is not used in MCNP6.3 to propagate response uncertainty without external analysis. Here, the TOol For Fast Error Estimation (TOFFEE) is a Python-based code developed to automate the propagation of cross-section uncertainty for MCNP evaluations. TOFFEE implements the sandwich rule to calculate the uncertainty from cross sections with sensitivity coefficients from MCNP6.3 and ENDF/B covariance data. In this paper, TOFFEE has been tested with benchmark experiments, and it has been compared to the uncertainty quantification capabilities of Sampler and TSUNAMI, within SCALE, to verify the application’s capabilities.

97 MATHEMATICS AND COMPUTING↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

A Novel Segmentation Algorithm for the ARM User Facility All-Sky Imagers Using Machine Learning Applications

Cloud cover plays a pivotal role in modulating the Earth's energy budget through the reflection of incoming solar radiation and the trapping of outgoing longwave radiation. Ground-based all-sky imagers offer an objective assessment of cloud cover that can be used to estimate solar irradiance, classify cloud types, track cloud movement, and serve as a benchmark 10 for the evaluation of satellite and reanalysis data products. The Atmospheric Radiation Measurement (ARM) user facility has utilized all-sky imagers for more than 25 years to monitor cloud cover and augment its comprehensive suite of atmospheric measurements. Following the retirement of its Total Sky Imager (TSI), ARM recently deployed the TSI’s successor, the All Sky Imager (ASI-16 camera systems). To provide a smooth transition and continuity to the vast amount of knowledge gathered by the TSI over the years, while addressing typical deployment issues, we developed a novel pixel segmentation algorithm, 15 the ASI Sky Cover (ASISKYCOVER). ASISKYCOVER builds on the different strengths and properties of the TSI processing algorithm while integrating machine learning techniques, ensuring data validity and accuracy across diverse atmospheric conditions. It enhances cloud cover characterization with new features such as artifact detection and uncertainty quantification. ASISKYCOVER also includes cloud cover estimates for near-zenith (narrow field-of-view) and reduces susceptibility to false detections. This study introduces ASISKYCOVER, details its algorithm framework, and demonstrates its capabilities using a 20 year-long dataset from the ARM Southern Great Plains site. Comparisons with co-located TSI data and other ARM measurements, such as zenith-pointing radars and lidars, are presented, underscoring the ASISKYCOVER’s potential to improve cloud cover analyses and data evaluation efforts, as well as to be integrated into higher-level data products that synergize instrument suites to generate new and insightful information

Silber, Israel↗

Physics vs structure: A systematic benchmark of learning strategies for multi-zone building thermal dynamics

Recent advances in physics-informed and data-driven machine learning promise improved thermal models for advanced building control, yet there is limited quantitative evidence on when added physics structure and architectural complexity are beneficial. Here, this work presents a systematic benchmark of five representative system identification methods for modeling multi-zone building thermal dynamics: linear state-space models, multi-layer perceptrons, neural state-space models, neural ordinary differential equations, and physically-consistent neural networks. The methods are evaluated across multiple data regimes and zone coupling strategies. Using a high-fidelity multi-zone commercial building emulator, we examine short-term and long-term prediction accuracy, computational efficiency, and ease of development. Our results reveal critical trade-offs between prediction performance, model complexity, and physical consistency. We demonstrate that decoupled, nonlinear black-box models consistently outperform coupled physics-constrained architectures in both predictive accuracy and out-of-distribution robustness in majority of the test cases for the building type considered in the study. Our findings quantify the cost of complexity in building thermal modeling and provide concrete, actionable, scenario-based guidelines for selecting model classes for control-oriented applications.

Building thermal modeling↗

Window Observables for Benchmarking Parton Distribution Functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel “window observables” that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-𝑥, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different window observables that can be defined within a region of 𝑥 where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

lattice QCD↗

Window observables for benchmarking parton distribution functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel ``window observables'' that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-x, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different ``window observables'' that can be defined within a region of x where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Bayesian inference of anisotropic 2D small-angle scattering from sparse measurement

Here, we present a Bayesian inference framework for reconstructing anisotropic two-dimensional small-angle scattering (2D SAS) patterns from sparse, noisy, or partially missing data. The method combines a symmetry-aware angular basis with radial Gaussian process priors to enable accurate, training-free interpolation and denoising. Computational benchmarks demonstrate reliable recovery of both isotropic and high-order anisotropic features under severe data reduction. Experimental validations on stretched polymers, sheared wormlike micelles, and carbon fibers show improved fidelity and resolution compared to raw measurements, achieving comparable accuracy with up to 50-fold fewer detected neutrons. This approach enables quantitative structural analysis under low-flux, time-limited, or single-shot conditions, extending the applicability of 2D SAS techniques to compact neutron sources and mechanically driven soft matter systems undergoing transient structural changes.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗