Search NASA⌕ Search

SEARCH · Search NASA

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Benchmarking Density Functional Theory Methods for Efficient Calculations of a Strongly Correlated Li 1– x Ni 1– y O 2−δ System

Transition metal oxides (TMOs), such as LiNiO 2 , are promising candidates for energy storage and electronic devices due to their unique electronic properties, exceptional physical and chemical characteristics, and ability to adopt multiple oxidation states. However, accurately predicting their properties using mean-field density functional theory (DFT) is challenging due to the presence of strongly correlated d-electrons and the complex interplay between their structural, electronic, and magnetic responses. These challenges are further exacerbated by the need to model defects, surfaces, and interfaces, which require computationally efficient, large-scale simulations. To address these issues, we carry out a benchmark study on the Li 1–x NiO 2 system, evaluating the performance of several popular functionals. Our findings demonstrate that combining SCAN functional relaxation with single-step HSE calculations provides a practical and scalable computational strategy. This approach balances accuracy and efficiency, enabling high-throughput simulations of strongly correlated TMOs and improved predictive modeling capability of TMOs for practical applications.

25 ENERGY STORAGE↗

Rates of Sea‐Level Rise Are Highly Sensitive to Ice Viscosity Parameters in Model Benchmarks

Glacier flow plays a major role in current and future rates of globally averaged sea-level rise. The viscosity of glacial ice, controlling the rate of flow, decreases as stress increases and is highly sensitive to the value of the stress exponent, $n$, in the constitutive equation for viscous flow. Glaciologists and climate modelers almost exclusively assume $n=3$ when modeling ice flow and projecting sea-level rise through forward modeling. However, recent work suggests that $n\approx 4$ better fits observations, prompting the question: How sensitive are projections of sea-level rise to the value of $n$? We use an established community ice flow model and standard benchmark experiments designed as an idealized representation of Pine Island Glacier, West Antarctica. While initializing an $n=3$ model to match observations of an $n=4$ ice sheet is possible, we find that incorrectly assuming $n=3$ when in fact $n=4$ dramatically underestimates rates of sea-level rise. The scale of this error grows nonlinearly with the magnitude of the climate forcing, acting to increase projection uncertainties. Additionally, we find that models often account for this stress-dependent rheology mismatch during model initialization in a way that masks this rheological effect in the short term while leaving model outputs vulnerable to larger biases in longer-term projections. Initializations to observations of Pine Island Glacier display similar rheology-mismatch fingerprints to our idealized example.

climate sensitivity↗

Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools

Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses, and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5213 previously published case reports using the Phenopacket Schema, the Human Phenotype Ontology and Mondo disease ontology. Prompts generated from each phenopacket were sent to seven LLMs, including four generalist models and three LLMs specialized for medical applications. The same phenopackets were used as input to a widely used diagnostic tool, Exomiser, in phenotype-only mode. The best LLM ranked the correct diagnosis first in 23.6% of cases, whereas Exomiser did so in 35.5% of cases. While the performance of LLMs for supporting differential diagnosis has been improving, it has not reached the level of commonly used traditional bioinformatics tools. Future research is needed to determine the best approach to incorporate LLMs into diagnostic pipelines.

Reese, Justin T. [Lawrence Berkeley National Labor↗

Downfolding from ab initio to interacting model Hamiltonians: comprehensive analysis and benchmarking of the DFT+cRPA approach

Abstract Model Hamiltonians are regularly derived from first principles to describe correlated matter. However, the standard methods for this contain a number of largely unexplored approximations. For a strongly correlated impurity model system, here we carefully compare a standard downfolding technique with the best possible ground-truth estimates for charge-neutral excited-state energies and wave functions using state-of-the-art first-principles many-body wave function approaches. To this end, we use the vanadocene molecule and analyze all downfolding aspects, including the Hamiltonian form, target basis, double-counting correction, and Coulomb interaction screening models. We find that the choice of target-space basis functions emerges as a key factor for the quality of the downfolded results, while orbital-dependent double-counting corrections diminish the quality. Background screening of the Coulomb interaction matrix elements primarily affects crystal-field excitations. Our benchmark uncovers the relative importance of each downfolding step and offers insights into the potential accuracy of minimal downfolded model Hamiltonians.

Chemistry↗

Increasing the hardness of posiform planting using random QUBOs for programmable quantum annealer benchmarking

Posiform planting is a method for constructing QUBO instances with a unique planted solution that can be tailored to arbitrary connectivity graphs. In this study we investigate making posiform planted QUBOs computationally harder by fusing many smaller random Ising models, whose global minimum is computed classically, with posiform planted QUBOs. The unique ground state of the resulting QUBO is the concatenation of (exactly one of) the ground states of each smaller problem. Our method generates QUBO instances that have a unique solution, are native to the hardware graph, and have tunable computational hardness. We use our QUBOs to benchmark three D-Wave quantum annealing processors (with 563–5627 qubits), and compare them against simulated annealing and Gurobi. Surprisingly, we find that the D-Wave ground state sampling success rate is not dependent on the glued random QUBO size, and that some QUBO classes are solved at high success rates at short annealing times on the Zephyr processors.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as p2r, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these miniapps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The 0E2 benchmarks for PWR UO 2 decay heat: an analysis from the NEA WPNCS

This paper presents the work performed in the subgroup 16 of the Working Party for Nuclear Criticality Safety (WPNCS) of the OECD Nuclear Energy Agency. The main goal was to define two decay heat benchmarks for Spent Nuclear Fuel (one pincell and one assembly), perform calculations and compare and analyze the results in light of existing calorimetric measurements. The selected case is the PWR UO2 assembly 0E2, irradiated at the Ringhals-3 reactor and measured at the Clab facility in Sweden. In total, 21 institutes worldwide participated to the exercise, leading to 55 calculated results (named C). It was found that the measured decay heat values (E) can be satisfactorily reproduced with two-dimensional assembly calculations, leading to an average C/E value of 0.99, with an uncertainty (or one standard deviation) of ±0.01.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmarking the exponential ansatz for the Holstein model

Polarons are quasiparticles formed as a result of lattice distortions induced by charge carriers. The single-electron Holstein model captures the fundamentals of single polaron physics. We examine the power of the exponential ansatz for the polaron ground-state wavefunction in its coupled cluster, canonical transformation, and (canonically transformed) perturbative variants across the parameter space of the Holstein model. Our benchmark serves to guide future developments of polaron wavefunctions beyond the single-electron Holstein model.

Chemistry↗

Benchmarking third-order cluster perturbation theory for electronically excited states

In this study, we investigate the reliability of cluster perturbation (CP) theory applied to the calculation of electronically excited states through a comprehensive benchmark. In CP theory, perturbative corrections are added to the properties of a parent excitation space, which converge toward the properties of a target excitation space. For the CPS(D-n) model, perturbative corrections through order n are added to the coupled cluster singles (CCS) excitation energies to target the coupled cluster singles and doubles (CCSD) excitation energies. Through a comparative analysis of excitation energy calculations across a diverse set of molecules and wavefunction methods, we present a comprehensive evaluation of the accuracy of the third-order CPS(D) model, CPS(D-3), in calculating excitation energies. Further, our findings demonstrate that CPS(D-3) is a reliable alternative to established methods, particularly CCSD, while systematically overestimating the excitation energies compared to high-level coupled cluster methods such as CC3. These results highlight the strengths and limitations of CPS(D-3), as well as the promising directions for its future development.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using multiple high-resolution datasets to benchmark the energy exascale earth system model (E3SM) for renewable resource assessment

The United States is accelerating its shift toward a renewable energy system. However, renewable resources, which harness energy from the Earth system, are susceptible to both present-day climate variability and future climate change. For example, variations in regional climate can alter renewable energy production patterns and site viability. The use of high-resolution climate model projections can therefore facilitate and may be critical to long-term planning of renewable energy investments. However, climate models must first be validated for renewable resource assessment. This research employs multiple high-spatiotemporal-resolution datasets to assess the capability of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 North American Regionally Refined Model (E3SMv2-NARRM) for predicting multi-year climatological values of solar and wind energy capacity factors in the continental U.S., with a focus on regional and seasonal variability. Present-day E3SMv2-NARRM simulations are compared with reported utility-scale production data obtained from the Energy Information Administration (EIA). In addition, E3SMv2-NARRM data are evaluated against non-climate benchmark models from the National Renewable Energy Laboratory, including the Wind Integration National Dataset Toolkit and the National Solar Radiation Database (NSRDB), as well as three wind energy datasets from PLUSWIND. Our analysis indicates that solar capacity factors from E3SM closely match those from the NSRDB dataset. However, both datasets tend to overestimate values by 10% in comparison to EIA data. Furthermore, biases in wind capacity factors within E3SM are notably pronounced in the West Coast regions, where the seasonal cycle diverges from EIA data.

Energy forecasting, Capacity factor, Renewable ene↗

Relativistic core–valence-separated equation-of-motion coupled-cluster singles and doubles method: Efficient implementation and benchmark calculations

An efficient implementation for the relativistic exact two-component core–valence-separated equation-of-motion coupled-cluster singles and doubles (X2C-CVS-EOM-CCSD) method is reported. The explicit exclusion of pure valence excitations in the EOM-CCSD excited-state eigenvalue equations significantly improves the efficiency for calculations of core-excited states. Benchmark relativistic CVS-EOM-CC calculations with systematic inclusion of relativistic, correlation, and basis-set effects are shown to provide highly accurate results for core ionized and excited states involving heavy atoms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Depletion Benchmark Analysis on a Lead Fast Reactor Using PyARC/OpenMC

PyARC is a user-friendly fast reactor analysis tool that automates multiphysics workflows using the “extended suite” of Argonne Reactor Computation (ARC) codes by providing a single common input for model definition, code execution, and output post-processing. A lead fast reactor (LFR) benchmark model is used to perform depletion calculations using the newly integrated OpenMC depletion capability in PyARC, building on previous analysis using the ARC codes through PyARC and Serpent. Results for core lifetime k-effective, shutdown decay heat, and end-of-life heavy-metal inventory are compared to verify the PyARC/OpenMC integration against the PyARC/ARC workflow and Serpent for depletion analysis of LFR designs. The results show satisfactory agreement among all three methods, with remaining discrepancies largely attributable to differences in nuclear data libraries and decay-chain modeling detail rather than to fundamental modeling limitations.

Kiesling, Kalin R.↗

A transient near to far field transformation method and verification benchmarking procedure

The numerical calculation of electromagnetic far fields in the time-domain requires a near to far field transformation (NTFF) method. While time-domain NTFF methods for popular finite-difference time-domain (FDTD) approaches are well established, there is little discourse on NTFF methods for finite-element time-domain (FETD) codes. Here, this work is concerned with the development of an NTFF method for the Empire FETD code, which utilizes curl and divergence conforming elements. This discretization presents a difficulty in obtaining the equivalent electric current for the NTFF. Straightforward finite element interpolation of the fields is shown to give poor accuracy. Alternative interpolation methods are recommended. An expanding magnetic quadrupole pulse benchmark problem, which is fully developed in the appendices, provides the basis for quantitative comparison.

FETD↗

Benchmarking machine learning interatomic potentials via phonon anharmonicity

Abstract Machine learning approaches have recently emerged as powerful tools to probe structure-property relationships in crystals and molecules. Specifically, machine learning interatomic potentials (MLIPs) can accurately reproduce first-principles data at a cost similar to that of conventional interatomic potential approaches. While MLIPs have been extensively tested across various classes of materials and molecules, a clear characterization of the anharmonic terms encoded in the MLIPs is lacking. Here, we benchmark popular MLIPs using the anharmonic vibrational Hamiltonian of ThO 2 in the fluorite crystal structure, which was constructed from density functional theory (DFT) using our highly accurate and efficient irreducible derivative methods. The anharmonic Hamiltonian was used to generate molecular dynamics (MD) trajectories, which were used to train three classes of MLIPs: Gaussian approximation potentials, artificial neural networks (ANN), and graph neural networks (GNN). The results were assessed by directly comparing phonons and their interactions, as well as phonon linewidths, phonon lineshifts, and thermal conductivity. The models were also trained on a DFT MD dataset, demonstrating good agreement up to fifth-order for the ANN and GNN. Our analysis demonstrates that MLIPs have great potential for accurately characterizing anharmonicity in materials systems at a fraction of the cost of conventional first principles-based approaches.

interatomic potentials↗

Benchmarking universal machine learning interatomic potentials for rapid analysis of inelastic neutron scattering data

The accurate calculation of phonons and vibrational spectra remains a significant challenge, requiring highly precise evaluations of interatomic forces. Traditional methods based on the quantum description of the electronic structure, while widely used, are computationally expensive and demand substantial expertise. Emerging universal machine learning interatomic potentials (uMLIPs) offer a transformative alternative by employing pre-trained neural network surrogates to predict interatomic forces directly from atomic coordinates. This approach dramatically reduces computation time and minimizes the need for technical knowledge. In this paper, we produce a phonon database comprising nearly 5000 inorganic crystals to benchmark the performance of several leading uMLIPs. We further assess these models in real-world applications by using them to analyze experimental inelastic neutron scattering data collected on a variety of materials. Through detailed comparisons, we identify the strengths and limitations of these uMLIPs, providing insights into their accuracy and suitability for fast calculations of phonons and related properties, as well as the potential for real-time interpretation of neutron scattering spectra. Our findings highlight how the rapid advancement of AI in science is revolutionizing experimental research and data analysis.

inelastic neutron scattering↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

Primeval very low-mass stars and brown dwarfs – VIII. The first age benchmark L subdwarf, a wide companion to a halo white dwarf

ABSTRACT We report the discovery of five white dwarf + ultracool dwarf systems identified as common proper motion wide binaries in the Gaia Catalogue of Nearby Stars. The discoveries include a white dwarf + L subdwarf binary, VVV 1256−62AB, a gravitationally bound system located 75.6$^{+1.9}_{-1.8}$ pc away with a projected separation of 1375$^{+35}_{-33}$ au. The primary is a cool DC white dwarf with a hydrogen dominated atmosphere, and has a total age of $10.5^{+3.3}_{-2.1}$ Gyr, based on white dwarf model fitting. The secondary is an L subdwarf with a metallicity of [M/H] = $-0.72^{+0.08}_{-0.10}$ (i.e. [Fe/H] = $-0.81\pm 0.10$) and $T_{\rm eff}$ = 2298$^{+45}_{-43}$ K based on atmospheric model fitting of its optical to near infrared spectrum, and likely has a mass just above the stellar/substellar boundary. The subsolar metallicity of the L subdwarf and the system’s total space velocity of 406 km s−1 indicates membership in the Galactic halo, and it has a flat eccentric Galactic orbit passing within 1 kpc of the centre of the Milky Way every $\sim$0.4 Gyr and extending to 15–31 kpc at apogal. VVV 1256−62B is the first L subdwarf to have a well-constrained age, making it an ideal benchmark of metal-poor ultracool dwarf atmospheres and evolution.

Zhang, Z. H. (ORCID:000000033047607X)↗