Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

SEED Platform for Building Performance Standards Implementation Guide (French Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the French translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (Arabic Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Arabic translation of NREL/FS-5500-90691.

benchmarking↗

SEED Platform for Building Performance Standards Implementation Guide (Mandarin Translation)

This guide provides an overview of the Standard Energy Efficiency Data (SEED) Platform. The SEED Platform developed by the U.S. Department of Energy (DOE) to provide a low-cost, user-friendly tool for jurisdictions to launch and manage energy benchmarking and Building Performance Standard (BPS) programs. It has been translated into Spanish. This is the Mandarin translation of NREL/FS-5500-90691.

benchmarking↗

HELIUM LEAK TEST MODELING OF A SPENT NUCLEAR FUEL CANISTER

The U.S. Department of Energy (DOE) is considering the development of one or more federal consolidated interim storage facilities (CISFs) to be used to store commercial spent nuclear fuel (SNF) at locations in the U.S. One of the first technical challenges of a CISF is performing an inspection of SNF canisters upon their receipt to confirm they can be placed into the CISF’s licensed storage configuration. The canister receipt inspection is critical to CISF site operations. The test is conceived as being a helium (He) leak check, intended to confirm that the confinement boundary of a SNF canister is intact. SNF canisters are filled with He when they are sealed, so detection of a He leak indicates that a through-wall flaw has occurred in the canister confinement boundary. Other measurements are planned to occur upon canister receipt in addition to the He leak check such as krypton-85 measurements, which would indicate confinement breaches of one or more fuel rods in addition to a breach of the SNF canister. However, the He leak check has been identified as one of such high importance and has such significant technical challenges that a full-scale demonstration is needed to confirm the He leak test’s viability and to assist in planning relative to its operational requirements. A modeling methodology for simulating the He detection test was developed to help inform the test plan and the design of the test vessels. To develop the modeling methodology a detailed computational fluid dynamics (CFD) benchmark model was constructed to compare against leak rate test data from a transportation package for radioactive material. This report is focused on modeling efforts to simulate the benchmark leak test.

Suffield, Sarah R.↗

Using multiple high-resolution datasets to benchmark the energy exascale earth system model (E3SM) for renewable resource assessment

The United States is accelerating its shift toward a renewable energy system. However, renewable resources, which harness energy from the Earth system, are susceptible to both present-day climate variability and future climate change. For example, variations in regional climate can alter renewable energy production patterns and site viability. The use of high-resolution climate model projections can therefore facilitate and may be critical to long-term planning of renewable energy investments. However, climate models must first be validated for renewable resource assessment. This research employs multiple high-spatiotemporal-resolution datasets to assess the capability of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 North American Regionally Refined Model (E3SMv2-NARRM) for predicting multi-year climatological values of solar and wind energy capacity factors in the continental U.S., with a focus on regional and seasonal variability. Present-day E3SMv2-NARRM simulations are compared with reported utility-scale production data obtained from the Energy Information Administration (EIA). In addition, E3SMv2-NARRM data are evaluated against non-climate benchmark models from the National Renewable Energy Laboratory, including the Wind Integration National Dataset Toolkit and the National Solar Radiation Database (NSRDB), as well as three wind energy datasets from PLUSWIND. Our analysis indicates that solar capacity factors from E3SM closely match those from the NSRDB dataset. However, both datasets tend to overestimate values by 10% in comparison to EIA data. Furthermore, biases in wind capacity factors within E3SM are notably pronounced in the West Coast regions, where the seasonal cycle diverges from EIA data.

Energy forecasting, Capacity factor, Renewable ene↗

ChatPORT: Fine-Tuned LLM for Easy Code {PORT}ing

Fine-tuning existing LLMs for specialized tasks has become a very attractive alternative due to its low cost and quick development cycle. With many pre-trained LLMs available, it is an increasingly complex task to choose the correct model as the starting point or base model. In this work we discuss ChatPORT - a specialized fine-tuned LLM geared towards providing correctly translated codes from one programming model to another. We evaluate a number of base models and compare and contrast their features and characteristics that make them a viable starting point. In this paper, we focus on the OpenMP offload porting capabilities of ChatPORT. We build our training data using kernels from the Heterogeneous Computing Benchmarks (HeCBench) [12] and the OpenMP Validation and Verification suite [5] to fine-tune the base models. We then test the model using unseen kernels extracted from the HeCBench benchmark suite. Our results show that: (1) not all open LLMs geared towards HPC are aware of programming models like OpenMP, (2) although all base models benefit from fine-tuning they learn differently and produce different correctness rates, (3) depending on the memory size and compute resource available, different base models can be used for fine-tuning without significantly affecting the quality of transpiled code they generate, (4) fine-tuning improved the correctness rate of the LLM by an average of 43.2%, and (5) feedback-based training data further increased the correctness rate by an average of 6% over the LLMs tested.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Benchmarking Variables for Checkpointing in HPC Applications

Checkpoint/Restart (C/R) is a widely used fault tolerance mechanism in converged systems of cloud, edge, and HPC. However, users often rely on their experience to determine which variables to checkpoint, as there is currently no benchmark that can provide a reference. This can result in checkpointing redundant or even incorrect variables. To address this issue, we propose a benchmark suite that includes critical variables for checkpointing, which have been manually identified, and a method for identifying those critical variables, with 20 representative HPC applications. Our method involves analyzing data dependency between variables to identify critical variables analytically. We verify the identified variables' correctness with a widely used C/R library FTI by an ablation study. With our benchmark suite and data dependency analysis, HPC practitioners now have a reference for identifying checkpointing variables and better knowledge of what kind of variables to checkpoint.

Fu, Xiang↗

Ecological Insights from Transferable Plant Biomass Mapping across the Arctic using High-resolution Structure-from-Motion and LiDAR Data

Warmer temperatures, permafrost thaw, and increased wildfire activity are driving rapid ecological change across the Arctic, significantly altering plant productivity and aboveground biomass (AGB). These rapid changes highlight the urgent need to improve monitoring of vegetation dynamics in the Earth’s northern ecosystems, where high spatiotemporal heterogeneity occurs at scales finer than those captured by traditional satellite observations. The growing use of Unoccupied Aerial Systems (UASs) presents an opportunity to overcome this limitation. Yet, the diversity of UAS platforms, sensors, and data collection and processing workflows presents challenges for developing standardized, generalizable approaches. To address this challenge, we compiled 672 AGB plots co-located with 183 UAS-based Structure-from-Motion (SfM) or Light Detection and Ranging (LiDAR) surveys collected across the Arctic. Here, we: (1) evaluated the generalizability of UAS-derived canopy structure derived from high-resolution SfM and LiDAR for estimating AGB, (2) assessed scaling errors and their sources in two recent satellite-based AGB products derived from Landsat and MODIS, and (3) demonstrated the use of high-resolution AGB maps to quantify biomass variation across tundra plant functional types (PFTs) and to monitor post-fire recovery. Our results show that both SfM and LiDAR accurately captured AGB and its variability across tundra PFTs using a Random Forest (RF) model (overall RMSE: 0.336 kg/m2), with mapping performance varying slightly by region and data source. Using UAS-derived AGB maps as a benchmark, we identified systematic biases in satellite-derived AGB products, largely attributable to the magnitude of AGB and structural heterogeneity within coarse-resolution pixels. Applying our model to repeat UAS surveys following a tundra fire on Seward Peninsula, we observed rapid AGB recovery in non-shrub patches, with biomass recovering to pre-fire levels within 2 years. In contrast, shrub patches recovered more slowly, with AGB gains continuing over 2–4 years through both in-patch growth and lateral expansion (via dispersal) into remaining burned areas. Overall, these findings demonstrate the generalizability of UAS-based SfM and LiDAR data for estimating tundra AGB and highlight the potential of our approach to be broadly applied to generate high-quality AGB data for ecological monitoring and model benchmarking across the Arctic.

Yang, Daryl [ORNL] (ORCID:0000000317057823)↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled datasets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed reinforcement learning model, PIRL-AD: Physics-Informed Reinforcement Learning for Anomaly Detection, that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a non-physics informed RL benchmark and a support vector machine (SVM) on data collected from a forced flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous datasets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% vs. 91.06% vs. 83.01%) and F1-score (respectively, 89.23% vs. 86.98% vs. 77.01%). Furthermore, the findings suggest the potential of physics-integrated reinforcement learning models for enhanced anomaly detection performance in NPPs.

Reinforcement learning↗

Towards robust surrogate models: Benchmarking machine learning approaches to expediting phase field simulations of brittle fracture

Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.

Benchmark dataset↗

Spectral line identification from a photoionised silicon plasma in emission

Next-generation X-ray satellite telescopes such as XRISM, NewAthena and Lynx will enable observations of exotic astrophysical sources at unprecedented spectral and spatial resolution. Proper interpretation of these data demands that the accuracy of the models is at least within the uncertainty of the observations. One set of quantities that might not currently meet this requirement is transition energies of various astrophysically relevant ions. Current databases are populated with many untested theoretical calculations. Accurate laboratory benchmarks are required to better understand the coming data. We obtained laboratory spectra of X-ray lines from a silicon plasma at an average spectral resolving power of ∼7500 with a spherically bent crystal spectrometer on the Z facility at Sandia National Laboratories. Many of the lines in the data are measured here for the first time. We report measurements of 53 transitions originating from the K-shells of He-like to B-like silicon in the energy range between ∼1795 and 1880 eV (6.6–6.9 Å). The lines were identified by qualitative comparison against a full synthetic spectrum calculated with ATOMIC. The average fractional uncertainty (uncertainty/energy) for all reported lines is ∼5.4 × 10 −5 . We compare the measured quantities against transition energies calculated with RATS and FAC as well as those reported in the NIST ASD and XSTAR’s uaDB. Average absolute differences relative to experimentally measured values are 0.20, 0.32, 0.17 and 0.38 eV, respectively. All calculations/databases show good agreement with the experimental values; NIST ASD shows the closest match overall.

astrophysical plasmas↗

Identifying Topological Defects in Lamellar Phases through Contour Analysis of Complex Wave Fields

Lamellar phases frequently contain structural imperfections that significantly affect their behaviors and properties. Our previous research successfully reconstructed real-space configurations of defective lamellar phases from diffuse scattering patterns, indicating the presence of phase vortices as a potential method for identifying topological defects disrupting the smectic ordering. Here, this report presents a mathematical framework using regularized wave fields to represent defective lamellar structures in real space. Phase singularities, resulting from the interference of random waves and indicating lamellar order disruption, are identified through a contour integral. These wave fields, derived from coherent scattering in reciprocal space, were validated via computational benchmarks analyzing small-angle neutron scattering data from AOT surfactant solutions, facilitating further statistical analysis of the defects. Our study highlights the potential to extract meaningful information about topological defects in lyotropic phases by inversely analyzing experimentally measured two-point static correlations. Our method allows for detailed structural analysis of various lyotropic phases, both particulate and nonparticulate, in their quiescent states and facilitates quantitative investigation of defects’ role in phase transitions. By integrating small-angle scattering, deep learning, and vortex tangle analysis, our comprehensive approach shows promise in addressing complex challenges in the structural analysis of soft matter systems.

36 MATERIALS SCIENCE↗

Complexity of many-body interactions in transition metals via machine-learned force fields from the TM23 data set

Abstract This work examines challenges associated with the accuracy of machine-learned force fields (MLFFs) for bulk solid and liquid phases ofd-block elements. In exhaustive detail, we contrast the performance of force, energy, and stress predictions across the transition metals for two leading MLFF models: a kernel-based atomic cluster expansion method implemented using sparse Gaussian processes (FLARE), and an equivariant message-passing neural network (NequIP). Early transition metals present higher relative errors and are more difficult to learn relative to late platinum- and coinage-group elements, and this trend persists across model architectures. Trends in complexity of interatomic interactions for different metals are revealed via comparison of the performance of representations with different many-body order and angular resolution. Using arguments based on perturbation theory on the occupied and unoccupieddstates near the Fermi level, we determine that the large, sharpddensity of states both above and below the Fermi level in early transition metals leads to a more complex, harder-to-learn potential energy surface for these metals. Increasing the fictitious electronic temperature (smearing) modifies the angular sensitivity of forces and makes the early transition metal forces easier to learn. This work illustrates challenges in capturing intricate properties of metallic bonding with current leading MLFFs and provides a reference data set for transition metals, aimed at benchmarking the accuracy and improving the development of emerging machine-learned approximations.

Chemistry↗

Accurately simulating core-collapse self-interacting dark matter halos

The properties of satellite halos provide a promising probe for dark matter (DM) physics. Observations have motivated current efforts to explain surprisingly compact DM halos. If DM is not collisionless, but has strong self-interactions, halos can undergo gravothermal collapse, leading to higher densities in the central region of the halo. However, it is challenging to model this collapse phase from first principles. To improve on this, we sought to better understand the numerical challenges and convergence properties of self-interacting dark matter (SIDM) N-body simulations in the collapse phase. Especially, our aim was to better understand the evolution of satellite halos. To do so, we ran SIDM N-body simulations of a low-mass halo in isolation and within an external gravitational potential. The simulation set-up was motivated by the perturber of the stellar stream GD-1. We find that the halo evolution is very sensitive to energy conservation errors, and a SIDM kernel size that is too large can artificially speed up the collapse. Moreover, we demonstrate that the King model can describe the density profile at small radii for the late stages that we have simulated. Furthermore, for our most highly resolved simulation (N = 5 × 10 7 ) we have made the data public. It can serve as a benchmark. Overall, we find that the current numerical methods do not suffer from convergence problems in the late collapse phase and provide guidance on how to choose numerical parameters, for example that the energy conservation error is better kept well below 1%. This allows simulations to be run of halos that become concentrated enough to explain observations of GD-1-like stellar streams or strong gravitational lensing systems.

dark matter↗

Characterizing laser-heated polymer foams with simultaneous x-ray fluorescence spectroscopy and Thomson scattering at the Matter in Extreme Conditions Endstation at LCLS

Understanding the behavior of polymer foams at high energy density conditions is crucial to advance inertial fusion energy research. Here, we present a new experimental platform designed to measure the thermodynamic state of these materials at megabar pressures. At the Matter in Extreme Conditions Endstation of the Linac Coherent Light Source, we heat samples using an optical, high-intensity, femtosecond laser and dynamically probe them with ultra-short, coherent x-ray pulses of high peak brightness. We perform x-ray Thomson scattering measurements in forward and backward scattering geometries to capture both collective and non-collective electron behavior in the sample. Simultaneously, x-ray fluorescence spectroscopy is used to measure the emission from a mid-Z dopant, providing complementary information on the plasma conditions. By combining these techniques, we obtain temporally resolved temperature measurements of the transient warm dense matter states. Our initial experiments designed to benchmark the platform with carbon samples yielded data resolving the ultrafast response to laser heating with sub-picosecond resolution, measuring plasma temperatures exceeding 50 eV. These findings lay the foundation for precision studies of the dynamic evolution of laser-heated polymer foams.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Bayesian Gaussian process inference for neutron spin echo measurement

Neutron spin echo (NSE) spectroscopy provides unique access to microscopic dynamics, but its application is often constrained by low neutron flux, long acquisition times, and significant noise. Here, we present a Bayesian inference approach based on Gaussian process regression (GPR) to reconstruct high-quality spin echo signals from sparse and noisy data by exploiting correlations in reciprocal space. Benchmarks on synthetic datasets and validation with experimental NSE measurements of dendrimers show that GPR suppresses noise, interpolates missing intensity values, and accommodates irregular observations. The method improves accuracy, shortens acquisition times, and enables high-throughput and real-time studies. Beyond NSE, the framework is broadly applicable to other low signal-to-noise ratio scattering techniques, thereby extending the scope of neutron spectroscopy.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗