Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Benchmarking machine learning interatomic potentials via phonon anharmonicity

Abstract Machine learning approaches have recently emerged as powerful tools to probe structure-property relationships in crystals and molecules. Specifically, machine learning interatomic potentials (MLIPs) can accurately reproduce first-principles data at a cost similar to that of conventional interatomic potential approaches. While MLIPs have been extensively tested across various classes of materials and molecules, a clear characterization of the anharmonic terms encoded in the MLIPs is lacking. Here, we benchmark popular MLIPs using the anharmonic vibrational Hamiltonian of ThO 2 in the fluorite crystal structure, which was constructed from density functional theory (DFT) using our highly accurate and efficient irreducible derivative methods. The anharmonic Hamiltonian was used to generate molecular dynamics (MD) trajectories, which were used to train three classes of MLIPs: Gaussian approximation potentials, artificial neural networks (ANN), and graph neural networks (GNN). The results were assessed by directly comparing phonons and their interactions, as well as phonon linewidths, phonon lineshifts, and thermal conductivity. The models were also trained on a DFT MD dataset, demonstrating good agreement up to fifth-order for the ANN and GNN. Our analysis demonstrates that MLIPs have great potential for accurately characterizing anharmonicity in materials systems at a fraction of the cost of conventional first principles-based approaches.

interatomic potentials↗

New U.S. Data Tools are Playing a Crucial Role in Decarbonizing Buildings at Speed, Scale, and Low Cost

Preparing buildings for retrofits traditionally requires expensive on-site audits or timeintensive simulation models. As a result, the majority of buildings fail to pursue cost-saving retrofits. To address these barriers, the U.S. Department of Energy (DOE) has introduced the Building Efficiency Targeting Tool for Energy Retrofits (BETTER)-a new, free, on-line tool that utilizes a data-driven analytical engine and user-friendly web interface to automatically analyze a building's monthly energy usage in response to weather conditions. The tool benchmarks a building's electric and fossil energy usage against peers; estimates energy, cost, and emissions reductions at the building and portfolio levels; recommends energy efficiency measures; and prioritizes buildings for net-zero energy retrofits. Thanks to interoperability with the DOE's Standard Energy Efficiency Data (SEED) platform, BETTER is supporting U.S. jurisdictions to prepare buildings for retrofit at speed, scale, and low cost to comply with energy policies. This paper discusses the use of BETTER and SEED by one of the branches of the California state government to streamline a retrofit program across 455 public non-residential buildings to align with state goals to reduce greenhouse gas emissions. It describes the organization's challenge to reduce energy consumption across a geographically diverse, aging portfolio; explores how BETTER and SEED improved workflow efficiency; presents preliminary results, including avoiding audit costs of $3.28 million and developing the groundwork for retrofit projects estimated to prevent emission of 2,271 t CO2e annually; and provides guidance for other jurisdictions seeking similar results.

BETTER↗

New U.S. Data Tools are Playing a Crucial Role in Decarbonizing Buildings at Speed, Scale, and Low Cost

Preparing buildings for retrofits traditionally requires expensive on-site audits or time- intensive simulation models. As a result, the majority of buildings fail to pursue cost-saving retrofits. To address these barriers, the U.S. Department of Energy (DOE) has introduced the Building Efficiency Targeting Tool for Energy Retrofits (BETTER)—a new, free, on-line tool that utilizes a data-driven analytical engine and user-friendly web interface to automatically analyze a building’s monthly energy usage in response to weather conditions. The tool benchmarks a building’s electric and fossil energy usage against peers; estimates energy, cost, and emissions reductions at the building and portfolio levels; recommends energy efficiency measures; and prioritizes buildings for net-zero energy retrofits. Thanks to interoperability with the DOE’s Standard Energy Efficiency Data (SEED) platform, BETTER is supporting U.S. jurisdictions to prepare buildings for retrofit at speed, scale, and low cost to comply with energy policies. This paper discusses the use of BETTER and SEED by one of the branches of the California state government to streamline a retrofit program across 455 public non-residential buildings to align with state goals to reduce greenhouse gas emissions. It describes the organization’s challenge to reduce energy consumption across a geographically diverse, aging portfolio; explores how BETTER and SEED improved workflow efficiency; presents preliminary results, including avoiding audit costs of $3.28 million and developing the groundwork for retrofit projects estimated to prevent emission of 2,271 t CO2e annually; and provides guidance for other jurisdictions seeking similar results.

Li, han↗

Data from: "Reply to ‘The challenge of defining effectively-no-snow’"

This repository contains the data and code associated with the paper titled "Reply to ‘The challenge of defining effectively-no-snow’" published in Nature Reviews Earth and Environment, 2026. In this reply, we argue that the 10th percentile of peak SWE (Snow Water Equivalent), which we propose in the original article, can be used as intended given it's a standardized, impact-based benchmark for comparing snow conditions across regions, not as a literal measure of snow absence. We present new evidence with SNOwpack TELemetry (SNOTEL) data showing that years meeting the threshold are overwhelmingly associated with subsequent drought (given United States Drought Monitor conditions), supporting its hydrologic and societal relevance. We conclude that while the distinction between "effectively no snow" and "zero snow" should be clearly communicated, the original definition remains appropriate for assessing impacts on snow-dependent water systems. The file code_nree_ML_reply_2026.Rmd contains the main processing scripts which analyze the SNOTEL data. Data from the US Drought Monitor was downloaded at: https://usdmdataservices using the Get Drought Severity Statistics By Area Percent' option, saved to the *_HUC4_delineated.csv files (Hydrologic Unit Code), which are labeled accordingly. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > TERRESTRIAL HYDROSPHERE > SNOW/ICE↗

The PARADIGM Project: Case Study in Balancing Experiment Uncertainty with Design simplicity

Accurate nuclear data are required for simulations of many applications including nuclear criticality safety. Actinide nuclear data at intermediate energies (from 1 to 100s of keV) are imprecise and inaccurate, because of scarce differential data, and an insufficient theory approach to capture the structures expected in the data to yield evaluated nuclear data, and lack of integral data for proper validation. This is a known deficiency but has proved challenging to address. More specifically, only 5% of integral experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) benchmark suite address intermediate energies (Fig. 1). Associated calculated effective multiplication factor, k eff , values for these experiments are far outside the experimental uncertainties and are 25× further from experiment than for fast energies. These differences could either stem from systematic biases in nuclear data, experiments or both. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly reduce (by more than tens of percent) the uncertainties of intermediate energy actinide nuclear data. The PARADIGM project designed and intends to execute LANSCE (Los Alamos Neutron Science CEnter) and NCERC (National Criticality Experiments Research Center) intermediate experiments in parallel. They will specifically address a high priority nuclear data need—reducing bias and uncertainty in intermediate plutonium nuclear data. The two experiment will achieve that by informing each other and nuclear theory. By doing all these steps in parallel, the timeline to deliver improved nuclear data to users will significantly be reduced. This work will focus on the integral experiment final design and the balance of design and modeling simplicity while minimizing experiment uncertainty.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SIDM Concerto: Compilation and Data Release of Self-interacting Dark Matter Zoom-in Simulations

We present SIDM Concerto: 14 cosmological zoom-in simulations in cold dark matter (CDM) and self-interacting dark matter (SIDM) models based on the Symphony and Milky Way-est suites. SIDM Concerto includes one Large Magellanic Cloud– (LMC-) mass system (host mass ∼10 11 M ⊙ ), two Milky Way (MW) analogs (∼10 12 M ⊙ ), two group-mass hosts (∼10 13 M ⊙ ), and one low-mass cluster (∼10 14 M ⊙ ). Each host contains ≈2 × 10 7 particles and is run in CDM and one or more strong, velocity-dependent SIDM models. Our analysis of SIDM (sub)halo populations over seven subhalo mass decades reveals that (1) the fraction of core-collapsed isolated halos and subhalos peaks at a maximum circular velocity corresponding to the transition of the SIDM cross section from a v −4 to v 0 scaling; (2) SIDM subhalo mass functions are suppressed by ≈50% relative to CDM in LMC, MW, and group-mass hosts but are consistent with CDM in the low-mass cluster host; (3) subhalos’ inner density profile slopes, which are more diverse in SIDM than in CDM, are sensitive to both the amplitude and shape of the SIDM cross section. Our simulations provide a benchmark for testing SIDM predictions with astrophysical observations of field and satellite galaxies, strong lensing systems, and stellar streams. Data products are publicly available at doi:10.5281/zenodo.14933624.

dark matter↗

Search for a third-generation leptoquark coupled to a τ lepton and a b quark through single, pair, and nonresonant production in proton-proton collisions at $ \sqrt{s} $ = 13 TeV

A search is presented for a third-generation leptoquark (LQ) coupled exclusively to a τ lepton and a b quark. The search is based on proton-proton collision data at a center-of-mass energy of 13 TeV recorded with the CMS detector, corresponding to an integrated luminosity of 138 fb$^{−1}$. Events with τ leptons and a varying number of jets originating from b quarks are considered, targeting the single and pair production of LQs, as well as nonresonant t-channel LQ exchange. An excess is observed in the data with respect to the background expectation in the combined analysis of all search regions. For a benchmark LQ mass of 2 TeV and an LQ-b-τ coupling strength of 2.5, the excess reaches a local significance of up to 2.8 standard deviations. Upper limits at the 95% confidence level are placed on the LQ production cross section in the LQ mass range 0.5–2.3 TeV, and up to 3 TeV for t-channel LQ exchange. Leptoquarks are excluded below masses of 1.22–1.88 TeV for different LQ models and varying coupling strengths up to 2.5. The study of nonresonant ττ production through t-channel LQ exchange allows lower limits on the LQ mass of up to 2.3 TeV to be obtained.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Models implemented in the methodological approach to design the initial STEP first wall contour

The official Spherical Tokamak for Energy Production mission aims to demonstrate the ability to generate net electricity from fusion with the STEP Prototype Power plant. One of the key technological and engineering challenges in fusion power plants is managing the loads on the first wall within acceptable limits. Therefore, the conceptual design development of the STEP Prototype Power plant needs to be based on load estimates derived using legitimate plasma physics assumptions through dynamic and flexible tools. The current design foresees the STEP main chamber first wall to withstand steady-state heat loads of up to ~1 MW/m 2 , excluding critical regions expected to receive higher heat loads such as the baffle regions approaching the divertors. These critical areas will require ad hoc assessments and will be designed with the presence of limiters. This article focuses on the models and methodology adopted for designing the 2-D poloidal contour of the STEP first wall, based on the anticipated charged particle and radiation heat loads during normal operation. Firstly, the models adopted for calculating the charged particle and radiation heat loads are introduced. The first model is validated through benchmarking against the particle tracing code SMARDDA, while the second model is verified by comparing it with data from the MAST-U experiment. Secondly, the model used to design the 2-D first wall contour according to the heat loads is explained. We acknowledge that this preliminary design stage assumes certain simplifications, notably an axisymmetric geometry, for computational efficiency and clarity in presentation. It is understood that subsequent design phases will address the complexities of real-world engineering, including non-axisymmetric effects, transient plasma scenarios, and the impact of disruptions on the first wall design. Finally, an automatic procedure based on these models is presented for defining the 2-D poloidal contour of the STEP first wall to minimize heat loads, taking into account the need to radiate most of the alpha-particle and auxiliary heating power. Here, by providing an overview of the models, methodology, and an automatic procedure, this paper contributes to the design process of the STEP first wall, addressing the engineering challenges associated with fusion power plant development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Direct Observations of Solute Dispersion in Rocks With Distinct Degree of Sub‐Micron Porosity

Abstract The transport of chemical species in rocks is affected by their structural heterogeneity to yield a wide spectrum of local solute concentrations. To quantify such imperfect mixing, advanced methodologies are needed that augment the traditional breakthrough curve analysis by probing solute concentration within the fluids locally. Here, we demonstrate the application of asynchronous, multimodality imaging by X‐ray computed tomography (XCT) and positron emission tomography (PET) to the study of passive tracer experiments in laboratory rock cores. The four‐dimensional concentration maps measured by PET reveal specific signatures of the transport process, which we have quantified using fundamental measures of mixing and spreading. We observe that the extent of solute spreading correlate strongly with the strength of subcore‐scale porosity heterogeneity measured by XCT, while dilution is enhanced in rocks containing substantial sub‐micron porosity. We observe that the analysis of different metrics is necessary, as they can differ in their sensitivity to the strength and forms of heterogeneity. The multimodality imaging approach is uniquely suited to probe the fundamental difference between spreading and mixing in heterogeneous media. We propose that when multi‐dimensional data is available, mixing and spreading can be independently quantified using the same metric. We also demonstrate that one‐dimensional transport models have limited predictive ability toward the internal evolution of the solute concentration, when the model is solely calibrated against the effluent breakthrough curves. The data set generated in this study can be used to build realistic digital rock models and to benchmark transport simulations that account deterministically for rock property heterogeneity.

Kurotori, Takeshi [Department of Chemical Engineer↗

Experimental shock Hugoniot and melt boundary temperatures of aluminum

Aluminum is a ubiquitous component in dynamic compression, pulsed power, and other high energy density physics studies. Its high-pressure behavior and phase diagram are extensively studied standards in shockwave physics. While theoretical calculations and multiphase equations of state have been benchmarked to velocity measurements of loading and unloading waves, pressure and density under shock, and other mechanical data, experimental temperature data under these conditions have not been reported. We conducted a series of experiments shocking and releasing aluminum 6061 and 1100 samples into lithium fluoride windows. We measured temperature at the sample–window interface under steady compression and subsequent isentropic release. These results allow us to constrain the temperature of the solid Hugoniot and the boundary between the liquid and face-centered cubic solid phases.

Hartsfield, Thomas Murray [Sandia National Laborat↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

ENnUI : Exemplar Navigator Using Inertial Sensors

SAND2025-04794O ENnUI: Exemplar Navigator Using Inertial Sensors helps users understand and compare different navigation algorithms that rely on inertial sensors. This reference library offers a collection of reference mechanization equations and methods for estimating the position and movement of vehicles over time. Rather than aiming for the highest precision or performance, ENnUI focuses on creating a user-friendly framework that allows users to benchmark their own algorithms against a standard set of tools. It is useful for post-processing data and real-time navigation solutions. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Walker II, Michael [Sandia National Lab. (SNL-CA),↗

Preliminary Results on Bayesian Inverse UQ for OECD/NEA WPNCS Subgroup 14 Benchmark Exercise for Error Recovery and Experimental Coverage

The Organization for Economic Cooperation and Development (OECD) Working Party on Nucelar Criticality Safety (WPNCS) has proposed a benchmark exercise representative of neutronic behavior in criticality experiments. Here, the goal is to develop confidence in data assimilation techniques used to adjust nuclear data. Participants are given synthetic experimental models with associated measured data and asked to estimate the model parameters given the model and measurements as well as provide predictions for separate application models. In this work, we performed data assimilation using Bayesian inverse Uncertainty Quantification (UQ) with machine learning surrogate models to produce posterior parameter distributions for the requested parameters and posterior predictive distributions for the requested responses. Several experimental models are shown to insufficiently inform the posterior parameter distributions for the applications involved. However, given sufficient experimental data, posterior parameter estimates yielded reduced uncertainty in the response predictions of interest while covering the experimental data.

Bayesian Inference↗

Benchmarking the Use of BPM Quadrupole Moments to Measure Emittance

For the PIP-II program, transverse emittance in the Fermilab Booster must remain well controlled at higher bunch intensities. 4-plate beam position monitors (BPMs) have a small but measurable quadrupole moment, making it possible to infer transverse emittance. By compositing many BPMs together, it becomes possible to improve the quality of the quadrupole signal. The Fermilab Booster BPM system has been used to measure these quadrupole moments in the past year and derive emittances from them. Recent benchmarks show that the derived BPM emittances show similar emittance evolution and value to IPM and Multiwire data. This approach can both supplement and complement existing non-intercepting emittance monitors in accelerators.

Balcewicz, M. A. [Fermilab]↗

Benchmarking the Use of BPM Quadrupole Moments to Measure Emittance

For the PIP-II program, transverse emittance in the Fermilab Booster must remain well controlled at higher bunch intensities. 4-plate beam position monitors (BPMs) have a small but measurable quadrupole moment, making it possible to infer transverse emittance. By compositing many BPMs together, it becomes possible to improve the quality of the quadrupole signal. The Fermilab Booster BPM system has been used to measure these quadrupole moments in the past year and derive emittances from them. Recent benchmarks show that the derived BPM emittances show similar emittance evolution and value to IPM and Multiwire data. This approach can both supplement and complement existing non-intercepting emittance monitors in accelerators.

Balcewicz, Michael A. [Fermilab] (ORCID:0009000557↗