Search NASA⌕ Search

SEARCH · Search NASA

Results for “application benchmarks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Application-level benchmarking of quantum computers using nonlocal game strategies

In a nonlocal game, two noncommunicating players cooperate to convince a referee that they possess a strategy that does not violate the rules of the game. Quantum strategies allow players to optimally win some games by performing joint measurements on a shared entangled state, but computing these strategies can be challenging. We present a variational quantum algorithm to compute quantum strategies for nonlocal games by encoding the rules of a nonlocal game into a Hamiltonian. We show how this algorithm can generate a short-depth optimal quantum strategy for a graph coloring game with a quantum advantage. This quantum strategy is then evaluated on fourteen different quantum hardware platforms to demonstrate its utility as a benchmark. Finally, we discuss potential sources of errors that can explain the observed decreased performance of the executed task and derive an expression for the number of samples required to accurately estimate the win rate in the presence of noise.

nonlocal games↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Quantum-Hardware Focused Application Performance Benchmarks (Final Technical Report)

Quantum computers promise to transform how we do scientific calculations. In this project, we benchmark different current quantum computers for solving chemistry problems and find ways to build noise-resilient implementations. We also use the characterization techniques developed to test ion trap quantum computers.

74 ATOMIC AND MOLECULAR PHYSICS↗

Data-driven analysis to understand GPU hardware resource usage of optimizations

With heterogeneous systems, the number of GPUs per chip increases to provide computational capabilities for solving science at a nanoscopic scale. However, low utilization for single GPUs defies the need to invest more money in expensive accelerators. Although related work develops optimizations to improve application performance, none studies how these optimizations impact hardware resource usage or average GPU utilization. Here, this paper takes a data-driven analysis approach in addressing this gap by (1) characterizing how hardware resource usage affects device utilization, execution time, or both, (2) presenting a multiobjective metric to identify important application-device interactions that can be optimized to improve device utilization and application performance jointly, (3) studying hardware resource usage behaviors of several optimizations for a benchmark application, and finally (4) identifying optimization opportunities for several scientific proxy applications based on their hardware resource usage behaviors. Furthermore, we demonstrate the applicability of our methodology by applying the identified optimizations to a proxy application, which improves the execution time, device utilization, and power consumption by up to 29.6%, 5.3% and 26.5% respectively.

Computer science↗

Participation in and Assessment of the Third DNCSH Public Workshop

Advanced reactor development and deployment is introducing fuel forms, materials, and configurations that differ from those used in traditional light-water reactors (LWRs), including tristructural-isotropic (TRISO)-based fuels, moderators, high-temperature materials, and compact core designs. These attributes can fall outside the range of applicability of existing criticality safety validation bases and may require additional work in code validation and benchmark applicability.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Quantum Computing for High-Energy Physics: State of the Art and Challenges

Quantum computers offer an intriguing path for a paradigmatic change of computing in the natural sciences and beyond, with the potential for achieving a so-called quantum advantage—namely, a significant (in some cases exponential) speedup of numerical simulations. The rapid development of hardware devices with various realizations of qubits enables the execution of small-scale but representative applications on quantum computers. In particular, the high-energy physics community plays a pivotal role in accessing the power of quantum computing, since the field is a driving source for challenging computational problems. This concerns, on the theoretical side, the exploration of models that are very hard or even impossible to address with classical techniques and, on the experimental side, the enormous data challenge of newly emerging experiments, such as the upgrade of the Large Hadron Collider. In this Roadmap paper, led by CERN, DESY, and IBM, we provide the status of high-energy physics quantum computations and give examples of theoretical and experimental target benchmark applications, which can be addressed in the near future. Having in mind hardware with about 100 qubits capable of executing several thousand two-qubit gates, where possible, we also provide resource estimates for the examples given using error-mitigated quantum computing. The ultimate declared goal of this task force is therefore to trigger further research in the high-energy physics community to develop interesting use cases for demonstrations on near-term quantum computers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Molecular simulation using transfer-learned potentials for the disordered nanoscale structure of nitrogen-doped nanoporous carbons

Machine learning (ML)-based molecular dynamics (MD) simulations of the formation of a class of N-doped nanoporous carbons are performed to assess their disordered partially graphitized nanoscale structure. The study is motivated by the effectiveness of so-called nitrogen assembly carbons (NACs) for catalysis applications. Benchmark simulations for pure-C disordered graphitic systems reveal the importance of reliably capturing the vdW component of the potentials in order to accurately describe the tendency for layering of disordered graphene-like sheets. In our modeling, this is achieved by a transfer learning strategy incorporating features of the energetics from the optB88-vdW DFT functional into potentials initially trained with a less expensive functional, thereby providing a superior description of the pure-C systems. Generation from MD simulations of realistic partially graphitized structures is significantly more challenging for N-doped versus for pure C systems. However, such structures are achieved by a tailored MD simulation protocol mimicking the experimental synthesis process and in particular incorporating an annealing and subsequent quenching stages. Simulated PXRD patterns effectively reproduce the features of experimental observations for NACs, including the appearance of a prominent but broad (002) peak at around 25, and the development of another weaker feature associated with in-layer ordering of mixed C-N graphene-like sheets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Correlated pair ansatz with a binary tree structure

We develop an efficient algorithm to implement the recently introduced binary tree state (BTS) ansatz on a classical computer. BTS allows a simple approximation to permanents arising from the computationally intractable antisymmetric product of interacting geminals and respects size-consistency. We show how to compute BTS overlap and reduced density matrices efficiently. We also explore two routes for developing correlated BTS approaches: Jastrow coupled cluster on BTS and linear combinations of BT states. Here, the resulting methods show great promise in benchmark applications to the reduced Bardeen–Cooper–Schrieffer Hamiltonian and the one-dimensional XXZ Heisenberg Hamiltonian.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Non-linear MHD modelling of transients in tokamaks: a review of recent advances with the JOREK code

Transient magneto-hydrodynamic (MHD) events like edge localized modes (ELMs) or disruptions are a concern for magnetic confinement fusion power plants. Research with the MHD code JOREK towards understanding control of such instabilities is reviewed here in a concise way to provide a complete overview, while we refer to the original publications for details. Experimental validation for unmitigated vertical displacement events progressed. The mechanism of vertical force mitigation by impurity injection was identified. Two-way eddy current coupling to CARIDDI was completed. Shattered pellet injection was simulated in JET, KSTAR, ASDEX Upgrade (AUG) and ITER. Benign runaway electron beam termination in JET and ITER was studied. Coupling of kinetic REs to the MHD is ongoing and a virtual RE synchrotron radiation diagnostic was developed. Regarding pedestal physics, regimes devoid of large ELMs in AUG were simulated and predictive JT60-SA simulations are ongoing. For ELM suppression by resonant magnetic perturbations (RMPs), AUG, ITER and EAST simulations were performed. A free boundary RMP model was validated against experiments. Evidence for penetrated magnetic islands at the pedestal top based on AUG experiments and simulations was found. Simulations of the naturally ELM-free quiescent H-mode in AUG and HL-3 show external kink mode formation prevents pedestal build-up towards an ELM within windows of the edge safety factor. With kinetic neutral particles, high field side high density formation in ITER was simulated and with kinetic impurities, tungsten transport in AUG RMP plasmas was studied. To capture turbulent transport, electro-static full-f particle in cell models for ion temperature gradient and trapped electron modes were established and benchmarked. Application to RMP plasmas shows enhanced turbulence in comparison to unperturbed states. Energetic particle interactions with MHD were studied. Flux pumping that prevents the safety factor on axis from dropping below unity was simulated. First non-linear stellarator applications include current relaxation in $l$ = 2 stellarators, while verification for advanced stellarators progresses.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Interplay of freeze-in and freeze-out: Lepton-flavored dark matter and muon colliders

We study a lepton-flavored dark matter model and its signatures at a future muon collider. We focus on the less-explored regime of feeble dark matter interactions, which suppresses the dangerous lepton-flavor-violating processes, gives rise to dark matter freeze-in production, and leads to long-lived particle signatures at colliders. We find that the interplay of dark matter freeze-in and its mediator freeze-out gives rise to an upper bound of around TeV scales on the dark matter mass. The signatures of this model depend on the lifetime of the mediator and can range from generic prompt decays to more exotic long-lived particle signals. In the prompt region, we calculate the signal yield, study useful kinematics cuts, and report tolerable systematics that would allow for a 5 σ discovery. In the long-lived region, we calculate the number of charged tracks and displaced lepton signals of our model in different parts of the detector and uncover kinematic features that can be used for background rejection. We show that, unlike in hadron colliders, multiple production channels contribute significantly, which leads to sharply distinct kinematics for electroweakly charged long-lived particle signals. Ultimately, the collider signatures of this lepton-flavored dark matter model are common among models of electroweak-charged new physics, rendering this model a useful and broadly applicable benchmark model for future muon collider studies that can help inform work on detector design and studies of systematics. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Surrogate models for linear response

Linear response theory is a well-established method in physics and chemistry for exploring excitations of many-body systems. In particular, the quasiparticle random-phase approximation (QRPA) provides a powerful microscopic framework by building excitations on top of the mean-field vacuum; however, its high computational cost limits model calibration and uncertainty quantification studies. Here, we present two complementary QRPA surrogate models and apply them to study response functions of finite nuclei. One is a reduced-order model that exploits the underlying QRPA structure, while the other utilizes the recently developed parametric matrix model algorithm to construct a map between the system’s Hamiltonian and observables. Our benchmark applications, the calculation of the electric dipole polarizability of 180 Yb and the 𝛽-decay half-life of 80 Ni, show that both emulators can achieve 0.1%–1% accuracy while offering a 6–7 orders of magnitude speedup compared to state-of-the-art QRPA solvers. These results demonstrate that the developed QRPA emulators are well positioned to enable Bayesian calibration and large-scale studies of computationally expensive physics models describing the properties of many-body systems.

Beta decay↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

MCNP® Code Version 6.3.1 Release Notes

The Monte Carlo N-Particle® (MCNP® ) code is a general-purpose, continuous-energy, generalized-geometry, time-dependent, radiation transport code developed by the MCNP development team. MCNP calculations provide predictive capabilities that can replace expensive or impossible-to-perform experiments. Specific application problems include simulations of experimental diagnostics, intrinsic radiation, radiation detection and measurement, criticality safety, nuclear threat reduction and response, radiation health protection, nuclear weapons effects, and nuclear forensics. This MCNP code, version 6.3.1, follows the MCNP6.3.0 version. Since the release of MCNP6.3.0, a variety of bug fixes and code enhancements have been completed for MCNP6.3.1. A few new features have also been added to this release to support both ongoing research and the release of the latest ENDF/B-VIII.1 nuclear data library. The MCNP code, version 6.3.1, theory and user input information is documented in MCNP® Code Version 6.3.1 Theory & User Manual, the build guidance for various platforms is documented in MCNP® Code Version 6.3.1 Build Guide, and the verification and validation testing for various application benchmark test suites is documented in MCNP® Code Version 6.3.1 Verification & Validation Testing.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Investigation into the Performance Benefits of Exposing Network Backpressure in UPC++ and GASNet-EX

This document is a brief summary of the research, and supporting development efforts, conducted by the project "Investigation into Improving Dynamic Adaptivity to System-Level Asynchrony in UPC++". We tested the hypothesis "The UPC++ and GASNet-EX runtimes can expose information from the network stack that enables applications to dynamically adapt to congestion, improving total throughput". We present experimental results from both a microbenchmark and an application benchmark that support this hypothesis.

97 MATHEMATICS AND COMPUTING↗