Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗

Using FIPD and OPTD to Benchmark Metallic Fuel Performance

This report serves as an introduction, tutorial, and benchmark specification for out-of-pile tests on metallic fuel. It introduces a new user to the EBR-II legacy fuel performance test program and the fast reactor fuel performance databases built to preserve the records. It then details the information stored in each database and how to find it. A benchmark specification is included for a small set of out-of-pile tests on U-10Zr fuel to function as a tutorial demonstrating how the legacy fuel performance data sets stored in the FIPD and OPTD databases can be used together to benchmark fuel performance models for steady-state and transient performance.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Adapting CLUTCH methodology to multigroup TSUNAMI-3D for eigenvalue sensitivity calculations

The sensitivity of the eigenvalue to uncertainties in nuclear data and its evaluation are important for nuclear criticality safety. TSUNAMI-3D sequences within the SCALE code system offer several options to the user community for calculating eigenvalue sensitivity coefficients with multigroup (MG) and continuous energy (CE) 3D transport capabilities. TSUNAMI-3D sequences implement the adjoint-based perturbation theory with MG KENO code, the Contributon Linked eigenvalue sensitivity/Uncertainty estimation via Track length importance CHaracterization (CLUTCH) method with CE KENO code, and the Iterated Fission Probability (IFP) method with CE KENO and Shift codes. Each method has benefits and limitations depending on the problem that is run. The work presented here aims to adapt the CLUTCH method, which enables the Contributon method's mesh-free, memory-efficient approach for calculating adjoint-weighted tallies for sensitivity calculations, to the MG TSUNAMI-3D sequence. This application would eliminate the explicit adjoint KENO calculation, as well as the memory-consuming mesh flux moment tallies required by the conventional MG TSUNAMI-3D. Smaller memory footprints in the CLUTCH methodology and relatively shorter runtimes in MG KENO transport can make MG TSUNAMI-3D a viable method for some complex problems. Moreover, this adaptation allows MG sensitivity calculations with Shift, ORNL's next-generation high-performance Monte Carlo transport code, which currently does not offer any sensitivity capabilities with MG particle transport simulations. Initial implementation of the new MG TSUNAMI-3D sequence and its preliminary results with a selected critical benchmark experiment in the Verified, Archived Library of Inputs and Data (VALID) are presented in this study.

KENO↗

Bringing HPE Slingshot 11 support to Open MPI

The Cray HPE Slingshot 11 network is used on the new exascale systems arriving at the U.S. Department of Energy (DoE) laboratories (e.g., Frontier, Aurora, Perlmutter). As such, the support of this network is an important capability to meet the needs of exascale applications. Here, this article highlights recent work to develop supporting infrastructure to enable Open MPI to efficiently support these new platforms. A key component of this effort involves development of a new Open Fabrics Interface (OFI) provider, LinkX. We discuss the design and development of enhancements that take advantage of the new Slingshot 11 network and AMD GPUs. We include performance data from tests on the Frontier supercomputer using synthetic communication benchmarks, and the vendor provided MPI as a baseline for comparison. The tests demonstrate full functionality of Open MPI on the system and initial results show favorable performance when compared to the highly tuned vendor implementation.

97 MATHEMATICS AND COMPUTING↗

A Detailed Reaction Mechanism for Thiosulfate Oxidation by Ozone in Aqueous Environments

The ozone oxidation, or ozonation, of thiosulfate is an important reaction for wastewater processing, where it is used for remediation of mining effluents, and for studying aerosol chemistry, where its fast reaction rate makes it an excellent model reaction. Although thiosulfate ozonation has been studied since the 1950s, challenges remain in developing a realistic reaction mechanism that can satisfactorily account for all observed products with a sequence of elementary reaction steps. Here, we present novel measurements using trapped microdroplets to study the pH-dependent thiosulfate ozonation kinetics. We detect known products and intermediates, including SO32-, SO42-, S3O62-, and S4O62-, establishing agreement with the literature. However, we identify S2O42- as a new reaction intermediate and find that the currently accepted mechanism does not directly explain observed pH effects. Thus, we develop a new mechanism, which incorporates S2O42- as an intermediate and uses elementary steps to explain the pH dependence of thiosulfate ozonation. The proposed mechanism is tested using a kinetic model benchmarked to the experiments presented here, then compared to literature data. We demonstrate good agreement between the proposed thiosulfate ozonation mechanism and experiments, suggesting that the insights in this paper can be leveraged in wastewater treatment and in understanding potential climate impacts.

Deal, Alexandra M↗

Validated Reactive Force Field Quantifies MXene Interfacial Properties, Mechanics, and Thermal Transport

MXenes combine rich surface chemistry, mechanical strength, and high conductivity for a multitude of emerging applications. Predictive modeling supports accelerated materials designs and has been limited by the absence of validated and transferable force fields. Here, we introduce an interpretable, reactive INTERFACE force field (IFF and IFF-R) for Ti 3 C 2 T x MXenes that is trained based on chemical knowledge and achieves quantitative agreement with experiments across lattice parameters (<0.5%), density (<0.2%), liquid contact angles, Raman spectra, and the in-plane elastic modulus (∼320 GPa). The models cover surface terminations from hydroxyl (−OH) to fluorine (−F) groups and are extensible to other chemistries. We introduce pH-resolved surface chemistry and identify dopamine adsorption mechanisms at MXene–aqueous interfaces supported by QCM-D and UV–Vis experiments. The data reveal coplanar and perpendicular binding modes and concentration-dependent multilayer assembly. We predict previously inaccessible properties, including termination-dependent cleavage energies, interlayer shear moduli and dynamic shear failure, nanoindentation and brittle fracture, anisotropic in-plane and out-of-plane thermal conductivities, including the role of defects. Agreement with available experimental data is consistently close and exceeds DFT accuracy across the benchmark properties examined. The IFF/IFF-R model is compatible with CHARMM, AMBER, OPLS, and CVFF force fields for simulations of MXenes with diverse surface terminations, electrolyte interfaces, biointerfaces, and polymer composites without additional parameters. Parameter sets, 3D models, and analysis scripts are provided for community use. The validated, reactive, and transferable IFF framework facilitates predictive design of MXene-based films, membranes, sensing interfaces, and composites.

MXene↗

Regularization by denoising diffusion models for solving inverse PDE problems with application to full waveform inversion

Partial differential equation (PDE)-governed inverse problems are fundamental across various scientific and engineering applications; yet they face significant challenges due to nonlinearity, ill-posedness, and sensitivity to noise. Here, we introduce a computational framework, regularization by denoising using diffusion models for partial differential equations (RED-DiffEq), by integrating physics-driven inversion and data-driven learning. RED-DiffEq leverages pretrained diffusion models as a regularization mechanism for PDE-governed inverse problems. We apply RED-DiffEq to solve the full waveform inversion problem in geophysics, a challenging seismic imaging technique that seeks to reconstruct high-resolution subsurface velocity models from seismic measurement data. Our method shows enhanced accuracy and robustness compared to benchmark methods. Additionally, it exhibits strong generalization and domain decomposition capacity, enabling the inversion of more complex velocity models with larger domains than those used in training the diffusion model. Our framework can also be directly applied to diverse PDE-governed inverse problems.

Shan, Siming [Yale University, New Haven, CT (Unit↗

High-Resolution Laser Spectroscopy on the Hyperfine Structure of 255Fm (𝑍=100)

We report on high-resolution laser spectroscopy of 255Fm (𝑇1/2=20 h), one of the heaviest nuclides available from reactor breeding. The hyperfine structures in two different atomic ground-state transitions at 398.4 nm and 398.2 nm were probed by in-source laser spectroscopy at the RISIKO mass separator in Mainz, using the perpendicularly illuminated laser ion source and trap (PI-LIST) high-resolution ion source. Experimental results were combined with hyperfine fields from various atomic ab initio calculations, in particular using multiconfiguration Dirac-Hartree-Fock theory, as implemented in grasp18. In this manner, the nuclear magnetic dipole and electric quadrupole moments were derived to be 𝜇=−0.75⁢(5) 𝜇N and 𝑄s=+5.84⁢(13) eb, respectively. The magnetic moment indicates occupation of the 𝜈⁢7/2⁢[613] Nilsson orbital, while the large quadrupole moment confirms strong, stable prolate deformation consistent with systematics in the heavy actinides. Comparisons with available expectation values from nuclear theory show good agreement, providing a stringent benchmark for the used theoretical models. These results revise earlier data and establish 255Fm as a reference isotope for future high-resolution studies.

Ezold, Julie [ORNL] (ORCID:0000000250550022)↗

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat↗

Wire-arc Additive Manufacturing Benchmark

This is the dataset associated with the 2022 SRP Additive Manufacturing Prediction Challenge, originally hosted on Github at https://github.com/SRP-AM/SRP_AM_Prediction_Challenge. The benchmark was designed for validating prediction for the temperature history, residual stress, and distortion of an additively manufactured metal part with relatively simple geometry. A calibration problem with the same as-built geometry is provided with measured quantities of interest; including temperature histories at selective locations, post-build residual stress at selective locations, and overall distortion measurements. The challenge problem is presented with a different build sequence (i.e. thermal history). In this dataset, we include the actual recorded calibration and challenge measurements, as well as benchmark template files for testing predictions without incorporating the challenge data. Supplementary files around the materials and setup are available for transparency and reproducibility.

Bachus, Nicholas [UC Davis, Davis, CA]↗

U.S. Average Corn Ethanol Production Baseline for 45Q Life Cycle Analysis Version 1.0 (2018-2022)

U.S. Average Benchmark life cycle inventory documentation for ethanol production for calendar years 2018 - 2022. The documentation includes a technical report describing data sources and statistical methods used and an Excel file showing calculations. The calculations performed are used in the 45Q benchmark json-ld dataset, compatible with the NETL CO2U LCA Guidance Toolkit. Summary impact results will be included in the NETL CO2U LCA Documentation Spreadsheet within the Toolkit. To access the model referenced in the report, please visit https://www.netl.doe.gov/energy-analysis/details?id=55ab21f1-9238-4b63-9a7a-ae1a4cfa7b05.

09 BIOMASS FUELS↗

Pile Oscillator Evaluation for Flattop [Poster]

Useful method for high-precision measurements of integral cross section in various samples. The availability of benchmark-quality effective delayed neutron fraction (ßeff) values for nuclear data and code validation remains limited.

MCNP↗

bmdrc: Python package for quantifying phenotypes from chemical exposures with benchmark dose modeling

Though chemical exposures are known to potentially have negative impacts on health, including contributing to chronic diseases such as cancer, the quantitative contribution of risk is not fully understood for every chemical. A commonly used approach to quantify levels of risk is to measure the proportion of organisms (such as a total number of zebrafish on a plate or mice in a cage) with abnormal behavioral responses or morphology at increasing concentrations of chemical exposure. A particular challenge with processing the proportional data from these assays is the appropriate estimation of chemical concentration levels that result in malformations or acute toxicity, as these values typically vary between experimental measurements. The recommended approach by the Environmental Protection Agency (EPA) is to fit benchmark dose curves with specific filters and model fitting steps, which are crucial to properly processing the proportional data. Several tools exist for the fitting of benchmark dose response curves, but none are standalone Python libraries built to process both morphological and behavioral data as proportions with all the EPA recommended filters, filter parameters, models, and model parameters. Thus, here we present the benchmark dose response curve (bmdrc) Python library, which was built to closely follow these EPA guidelines with helpful visualizations of filters and fitted model curves, and reports for reproducibility purposes. bmdrc is open-source and has demonstrated utility as a support package to an existing web portal for information on chemicals (https://srp.pnnl.gov). Our package will support any toxicology analysis where the response is a proportional value at increasing levels of a concentration of a chemical or chemical mixture.

Superfund↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Porosity in nuclear graphite and its impact on nuclear reactor science and criticality safety applications

Porosity in nuclear-grade graphite significantly influences its low-energy neutron scattering, yet its effect on underlying phonon properties remains debated. Here, this work integrates inelastic and small-angle neutron scattering (INS/SANS) experiments, advanced atomistic simulations with a novel machine-learned potential (DeepMD), total cross-section measurements, and neutronics calculations (SCALE, MCNP, OpenMC) to investigate porosity’s impact on neutron thermalization. INS measurements on diverse graphite grades reveal no discernible porosity effect on phonon spectra, which align with crystalline graphite. Conversely, total cross-section data below ≈10 meV show increased scattering attributable to SANS. Our DeepMD simulations demonstrate that realistic micropores do not distort phonon spectra, challenging the assumptions in current ENDF/B-VIII.1 porosity thermal scattering laws (TSLs). These TSLs, based on random atom removal, produce unphysical phonon spectra and inflate inelastic cross-sections. Augmenting a crystalline TSL with an SANS component accurately captures experimental total cross-sections. Neutronics benchmarks (ICSBEP/IRPhE) show ENDF porosity TSLs unphysically increase neutron multiplication factor, keff. Crucially, incorporating SANS physics (NCrystal/OpenMC) indicates accurately modeled porosity negligibly affects keff, reactor physics, or criticality safety.

Critical benchmarks↗

NEA HTTR LOFC Project Test#3 Benchmark Results

In the second half of FY23, the High Temperature Engineering Test Reactor (HTTR) Loss Of Forced Cooling (LOFC)#3 data for the 9 MW test case with Vessel Cooling System (VCS) off were made available through the Nuclear Energy Agency (NEA) LOFC project framework; the neutronic model developed for the initial test (LOFC#1) achieved a satisfactory level of maturity, demonstrating its accuracy in predicting power evolution and core re-criticality, but LOFC#3 should be used primarily to investigate thermal hydraulic phenomena, as the reactor was shut down prematurely due to overheating in the upper reactor components, which prevented re-criticality; this report focuses on advancing the HTTR thermal hydraulic model to accurately simulate the LOFC#3 scenario, including simulating the LOFC#3 benchmark and generating the corresponding benchmark specifications, aiming to ensure consistency across participant models and provide essential data for future participants, including private industry stakeholders seeking to validate their computational tools.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

U.S. Average Nitrogen Fertilizer Production Baseline Documentation For 45Q Life Cycle Analysis: Version 1.0 (2018-2022)

U.S. Average Benchmark life cycle inventory documentation for nitrogenous fertilizer production for calendar years 2018 - 2022. The documentation includes a technical report describing data sources and statistical methods used and an Excel file showing calculations. The calculations performed are used in the 45Q benchmark json-ld dataset, compatible with the NETL CO2U LCA Guidance Toolkit. Summary impact results will be included in the NETL CO2U LCA Documentation Spreadsheet within the Toolkit. https://www.netl.doe.gov/energy-analysis/details?id=7c0db05a-231e-40a5-b5fd-112a7c001834

09 BIOMASS FUELS↗