Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmarks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Phonon Olympics: Phonon property and lattice thermal conductivity benchmarking from open-source packages

Three widely used open-source packages for determining phonon properties and lattice thermal conductivities (ALAMODE, phono3py, and ShengBTE) are benchmarked by teams of expert users and the package developers. The phonons for Ge, RbBr, monolayer MoSe 2 , and AlN are modeled at zero temperature, and they scatter through three-phonon and phonon-isotope processes, with thermal conductivities obtained from the linearized Peierls–Boltzmann transport equation with input from density functional theory calculations. Over a wide range of temperatures, the thermal conductivities calculated by the teams fall within at most ±15% of their mean values for each of the four materials. The phonon frequencies, obtained from the harmonic force constants, do not show large differences between the calculations, indicating that the modal heat capacities and group velocities are not responsible for the thermal conductivity variations. It is the lifetimes associated with three-phonon scattering, obtained from the cubic force constants, that drive the variations. The many decisions required to calculate the cubic force constants (e.g., supercell size, atomic displacement, neighbor cutoff, and application of symmetries) make identification of the precise origin of the thermal conductivity variations challenging. The calculated thermal conductivities do not generally show agreement with experimental measurements, which is attributed to the limitations of the density functional theory calculations. Guidance for the development of best practices is provided, which will help to standardize protocols needed for building thermal conductivity databases. The results provide a baseline for future benchmarking of other packages and more advanced calculations.

McGaughey, Alan J. H. [Carnegie Mellon Univ., Pitt↗

Benchmarking core turbulence and transport predictions for an inductive compact tokamak reactor plasma

Motivated by the need for accurate, timely, and efficient calculations of plasma transport, predictions of plasma turbulence properties made using different TGLF saturation rules are benchmarked against corresponding predictions from linear and nonlinear gyrokinetic CGYRO simulations. This benchmarking is carried out using parameters taken from an inductive burning plasma scenario in a hypothetical compact high-field (R maj = 4 m, B T = 8 T) tokamak, lying in a much different regime of parameter space than either the TGLF calibration regime or current-day experiments. The core turbulent transport in this scenario is predicted to be dominated by ion temperature gradient (ITG) turbulence. In general, the ITG critical gradients predicted by various TGLF saturation rules are quite close to the CGYRO predictions. Both codes predict similar linear ITG growth rates and frequency spectra, as well as their scaling with R/L T i = −Rd ln(T i )/dr. However, TGLF systematically predicts unstable trapped-electron modes (TEMs) above k y ρ s ≃ 0.5 not seen by CGYRO for the same parameters, due to TGLF predicting a lower threshold in R/L T e than CGYRO for TEM onset. It is shown that for this scenario, nonlinear CGYRO simulations predict stiffer ITG turbulence than the TGLF SAT0 and SAT1 saturation rules, with energy fluxes close in magnitude and scaling with R/L T i to what is predicted by the SAT2 saturation rule. Self-consistent core profiles calculated using nonlinear CGYRO flux predictions and the PORTALS transport solver are shown to agree fairly well with corresponding predictions made using the TGLF SAT2 model, including a similar level of density peaking.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Hofstadter butterfly and quantum transport benchmarks in PVA-exfoliated graphene heterostructures

Polymer exposure during van der Waals heterostructure fabrication is widely regarded as compromising the integrity of the electronic system required for hosting emergent quantum physics. This assumption has persisted largely because electronic benchmarking of heterostructures from polymer-exposed graphene has remained limited to only foundational transport metrics, such as mobility and charge inhomogeneity. Here, we challenge this assumption by establishing that graphene heterostructures produced by polyvinyl alcohol (PVA)-assisted exfoliation and encapsulated in hexagonal boron nitride using elevated-temperature lamination satisfy demanding quantum transport benchmarks. Beyond exhibiting ultra-high mobility and ballistic transport, these heterostructures yield quantum scattering times comparable to the best polymer-free devices. Most demanding of all, moiré superlattices from PVA-exposed graphene exhibit Hofstadter butterfly spectra, confirming spatially uniform interlayer coupling across the device area. These results establish that PVA exposure is compatible with low-disorder electronic systems, relaxing the trade-off between scalable fabrication and low-disorder quantum transport. This study motivates further development of polymer-assisted assembly with engineered residue-removal protocols.

36 MATERIALS SCIENCE↗

Role of electron correlation on the adenine dimer interaction for non-equilibrium geometries: a benchmark Quantum Monte Carlo study

The accurate description of non-covalent interactions is critical for understanding the structure, dynamics, and eventual function of biomolecules. The adenine dimer serves as a benchmark system for computational methods due to its role in nucleic acid structures and its rich conformational landscape. In this study, we employ benchmark diffusion quantum Monte Carlo (DMC) methods to investigate the relative energies and role of electron correlation on a set of adenine dimer conformations generated via a search of the potential energy landscape using the global optimizer algorithm. Relative DMC energies are compared against a wide range of density functional theory (DFT) approximation results. We find that although most of the DFT functionals perform well for low-energy structures, their accuracy varies significantly for higher-energy conformations, including stacked and T-shaped structures. A large fraction of the variation is due to the treatment of the van der Waals interaction. BLYP, B3LYP, and PBE0 significantly improve with added D4 dispersion, while the recent r2SCAN-D4 and ωB97M-V functionals show the least scatter and closest agreement with the DMC. These findings highlight the delicate nature of these interactions in biomolecular systems and provide guidance for simulations of their structure and dynamics and for the development of machine learned interatomic potentials.

Washburn, Laurel [ORNL] (ORCID:0000000324179335)↗

The Role of Nuclear Data Sensitivities in Prompt α-Eigenvalue Predictions of Delayed Critical Benchmarks

Alpha (α) eigenvalues, which describe the logarithmic time derivative of the neutron population in a multiplying system, are integral to time-dependent behavior and diagnostic applications. However, uncertainties in the evaluated nuclear data can significantly impact the accuracy of transport simulations for such quantities. This work explores the use of machine learning models to predict two key outputs, α-eigenvalues and keff bias, using input features derived from α-eigenvalue sensitivities to nuclear data. The criticality safety benchmark models used in this study come from the International Handbook of Evaluated Criticality Safety Benchmark Experiments. Three models, random forest, XGBoost, and NGBoost, are trained on both energy-resolved and energy-summed α sensitivities. For the α-eigenvalue bias prediction, NGBoost achieved the highest R 2 (0.9476) using energy-resolved features, while XGBoost performed best using summed sensitivities. In contrast, when predicting the keff bias, all the models showed moderate predictive capability (best R 2 ≈ 0.72), as the mapping from the static α-sensitivities to the static keff bias was less direct. SHAP (SHapley Additive exPlanations) analysis was used to interpret the model predictions. Across both prediction tasks, the features associated with neutron capture [H-1 (n, γ)], uranium scattering reactions (such as 235 U elastic/inelastic), and actinide capture/fission reactions (such as 239 Pu and 234 U) were consistently identified as the most impactful. This highlights the key role of specific nuclear reactions and energy ranges in shaping both time-dependent and steady-state criticality behavior. These results demonstrated that α-sensitivities, despite being computed for time-dependent metrics, can provide valuable insights for predicting both α-eigenvalues and the keff bias. Moreover, machine learning models offer a promising pathway for uncovering important nuclear data dependencies and guiding future data evaluation efforts.

Nuclear data↗

Benchmarking machine learning strategies for phase-field problems

Abstract We present a comprehensive benchmarking framework for evaluating machine-learning approaches applied to phase-field problems. This framework focuses on four key analysis areas crucial for assessing the performance of such approaches in a systematic and structured way. Firstly, interpolation tasks are examined to identify trends in prediction accuracy and accumulation of error over simulation time. Secondly, extrapolation tasks are also evaluated according to the same metrics. Thirdly, the relationship between model performance and data requirements is investigated to understand the impact on predictions and robustness of these approaches. Finally, systematic errors are analyzed to identify specific events or inadvertent rare events triggering high errors. Quantitative metrics evaluating the local and global description of the microstructure evolution, along with other scalar metrics representative of phase-field problems, are used across these four analysis areas. This benchmarking framework provides a path to evaluate the effectiveness and limitations of machine-learning strategies applied to phase-field problems, ultimately facilitating their practical application.

36 MATERIALS SCIENCE↗

Real singlet scalar benchmarks in the multi-TeV resonance regime

Scalar extensions of the Standard Model (SM) are of much interest at the Large Hadron Collider (LHC) and future colliders. In particular, these models can give rise to resonant di-Higgs production and alter the Higgs trilinear coupling. In this paper, we study di-Higgs production in the Standard Model extended by a real scalar singlet with no additional symmetries. We determine how large the resonant di-Higgs rate and variation in the Higgs trilinear coupling can be in four scenarios: current LHC results and projected results at the high luminosity LHC (HL-LHC), the HL-LHC combined with a circular 𝑒 − ⁢𝑒 + collider such as the Circular Electron Positron Collider or Future Circular Collider with electron-positron collisions, and the HL-LHC combined with a linear 𝑒 − ⁢𝑒 + collider such as the International Linear Collider. While these are updated results from a previous study by [I. M. Lewis and M. Sullivan, Benchmarks for double Higgs production in the singlet extended standard model at the LHC, Phys. Rev. D 96, 035037 (2017).] using current LHC data, we go further and find benchmark points in the multi-TeV resonance regime for future colliders beyond the HL-LHC. Considering current LHC results, the resonant di-Higgs rate can still be an order of magnitude larger than the SM predicted di-Higgs rate. In the HL-LHC scenario, the Higgs trilinear coupling can still be a factor of three larger than the SM prediction for resonance masses in the 1.5–3.5 TeV range, where resonant searches may have less reach. This enhancement is just at the projected 2⁢𝜎 sensitivity of the HL-LHC. We find there are resonance masses for which the change in the Higgs trilinear is maximized while the resonant rate is negligible. We provide an analytical understanding of these effects with a discussion on the interplay of various constraints on the parameter space and the Higgs trilinear coupling.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

Benchmark High-Fidelity EMT Models for Power Grid with PV Plants

In recent times electromagnetic transient (EMT) modeling tools have been identified as one of the most important requirements in replicating, analyzing, and investigating the dynamics of the power grid with photovoltaic (PV) plants. However, there are no benchmark models for power grid with PVs to investigate emerging challenges with higher penetration of PVs (like trips and momentary cessations during faults from a region far away). To this end, in this paper, synthetic benchmark high-fidelity EMT dynamic models of power grid with large-scale PV plants are presented. The models are developed in PSCAD and PSCAD/Fortran. Simulation results for different use cases (events) and scenarios are presented.

Marthi, Phani Ratna Vanamali↗

Lifetimes of excited states in $^{16}$C as a benchmark for ab initio developments

Lifetimes of higher-lying states ($2_2^+$ and $4_1^+$) in 16 C have been measured, employing the Gammasphere and Microball detector arrays, as key observables to test and refine ab initio calculations based on interactions developed within chiral Effective Field Theory. The presented experimental constraints to these lifetimes of $\tau ({2_2^+}) = [244, 446]\,~\textrm{fs}$ and $\tau ({4_1^+}) = [1.8, 4]\,~\textrm{ps}$, combined with previous results on the lifetime of the $2_1^+$ state of 16 C, provide a rather complete set of key observables to benchmark the theoretical developments. We present No-Core Shell-Model calculations using state-of-the-art chiral 2- (NN) and 3-nucleon (3N) interactions at next-to-next-to-next-to-leading order for both the NN and the 3N contributions and a generalized natural-orbital basis (instead of the conventional harmonic-oscillator single-particle basis) which reproduce, for the first time, the experimental findings remarkably well. The level of agreement of the new calculations as compared to the CD-Bonn meson-exchange NN interaction is notable and presents a critical benchmark for theory.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Reference solutions for linear radiation transport: the Hohlraum and Lattice Benchmarks

Radiation transport describes the propagation of energetic particles through space as they interact with a surrounding material medium. In a kinetic description, radiation transport is modeled by a radiation transport equation (RTE) that prescribes the density of the radiation in position-momentum phase space. The purpose of this dataset is to provide highly resolved solutions to two benchmark problems. These two benchmarks do not possess exact solutions; moreover, the construction of a manufactured solution may require a non-physical source that is not desirable, especially if it spoils the physical nature of the solution. Thus the goal of this computational study is to provide a highly resolved reference solution for testing newer, more cost efficient methods that are currently being developed in the research community.

97 MATHEMATICS AND COMPUTING↗

Li1−xNiO2 Many-body DMC Benchmark Dataset

The dataset contains all numerical data generated in support of the manuscript “Many‑body Benchmark of Electronic Charge and Spin Densities for Li1–xNiO2​” (Journal of Chemical Theory and Computation, DOI: 10.1021/acs.jctc.5c02097, URL: https://pubs.acs.org/doi/10.1021/acs.jctc.5c02097). The materials included in this repository are: 1. Data files used to produce all figures and tables in the main manuscript and supporting information. 2. Benchmark density‑functional theory (DFT) datasets used for the charge‑ and spin‑density analyses. 3. Reference many‑body diffusion Monte Carlo (DMC) calculations and associated input/output files.

36 MATERIALS SCIENCE↗

Wire-arc Additive Manufacturing Benchmark

This is the dataset associated with the 2022 SRP Additive Manufacturing Prediction Challenge, originally hosted on Github at https://github.com/SRP-AM/SRP_AM_Prediction_Challenge. The benchmark was designed for validating prediction for the temperature history, residual stress, and distortion of an additively manufactured metal part with relatively simple geometry. A calibration problem with the same as-built geometry is provided with measured quantities of interest; including temperature histories at selective locations, post-build residual stress at selective locations, and overall distortion measurements. The challenge problem is presented with a different build sequence (i.e. thermal history). In this dataset, we include the actual recorded calibration and challenge measurements, as well as benchmark template files for testing predictions without incorporating the challenge data. Supplementary files around the materials and setup are available for transparency and reproducibility.

Bachus, Nicholas [UC Davis, Davis, CA]↗

Studying the Random Number Generators in MCNP6 using an Analytic Benchmark

An analytic solution to a previously studied toy problem is derived and used as a code verification benchmark. Using various Random Number Generators (RNGs) in MCNP6, including the newest SFC64 RNG available in MCNP6.3.1, and their various properties (e.g., RNG stride), we show how these RNGs perform and how to correct or workaround potential issues with respect to the analytic benchmark problem.

97 MATHEMATICS AND COMPUTING↗

Modeling Enhancements, Cross-Section Generation Updates, and Benchmarking with Shift

This technical report documents the modeling enhancements, cross-section generation updates, and bench marking with the Shift Monte Carlo code performed under the US Department of Energy Nuclear Energy Advanced Modeling and Simulation Program in FY 2024. The work performed included several modeling enhancements, such as integration of cross-section generation in Titan and the ability to produce microscopic multigroup cross sections with Shift. Benchmarking of the cross sections produced by Shift and the two-step workflow with Griffin was performed for three problems: the Advanced Breeder Test Reactor, a generic pebble bed reactor, and a TRISO heat pipe microreactor. Comparisons of results from these benchmark problems were done with Serpent, OpenMC, and Griffin. These enhancements provide a robust foundation for applying Shift for both reference and two-step neutronics analysis for advanced reactor simulation.

97 MATHEMATICS AND COMPUTING↗