Search NASA⌕ Search

SEARCH · Search NASA

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Toward Well-Provenanced Computer System Benchmarking: An Update

This report provides an update on the current design, implementation, and usage of bueno, a Python framework enabled by container technology that supports gradations of reproducibility for wellprovenanced benchmarking of sequential and parallel programs. The ultimate goal of the bueno project is to provide convenient access to mechanisms that aid in the automated generation, collection, and dissemination of data relevant for experimental reproducibility in computer system benchmarking.

97 MATHEMATICS AND COMPUTING↗

The neurobench framework for benchmarking neuromorphic computing algorithms and systems

Neuromorphic computing shows promise for advancing computing efficiency and capabilities of AI applications using brain-inspired principles. However, the neuromorphic research field currently lacks standardized benchmarks, making it difficult to accurately measure technological advancements, compare performance with conventional methods, and identify promising future research directions. This article presents NeuroBench, a benchmark framework for neuromorphic algorithms and systems, which is collaboratively designed from an open community of researchers across industry and academia. NeuroBench introduces a common set of tools and systematic methodology for inclusive benchmark measurement, delivering an objective reference framework for quantifying neuromorphic approaches in both hardware-independent and hardware-dependent settings. For latest project updates, visit the project website (neurobench.ai).

Yik, Jason [Harvard Univ., Cambridge, MA (United S↗

QASMBench: A Low-Level Quantum Benchmark Suite for NISQ Evaluation and Simulation

The rapid development of quantum computing (QC) in the NISQ era urgently demands a low-level benchmark suite and insightful evaluation metrics for characterizing the properties of prototype NISQ devices, the efficiency of QC programming compilers, schedulers and assemblers, and the capability of quantum system simulators in a classical computer. In this work, we fill this gap by proposing a low-level, easy-to-use benchmark suite called QASMBench based on the OpenQASM assembly representation. It consolidates commonly used quantum routines and kernels from a variety of domains including chemistry, simulation, linear algebra, searching, optimization, arithmetic, machine learning, fault tolerance, cryptography, and so on, trading-off between generality and usability. To analyze these kernels in terms of NISQ device execution, in addition to circuit width and depth, we propose four circuit metrics including gate density, retention lifespan, measurement density, and entanglement variance, to extract more insights about the execution efficiency, the susceptibility to NISQ error, and the potential gain from machine-specific optimizations. Applications in QASMBench can be launched and verified on several NISQ platforms, including IBM-Q, Rigetti, IonQ and Quantinuum. For evaluation, we measure the execution fidelity of a subset of QASMBench applications on 12 IBM-Q machines through density matrix state tomography, comprising 25K circuit evaluations. In addition we also compare the fidelity of executions among the IBM-Q machines, the IonQ QPU and the Rigetti Aspen M-1 system.

97 MATHEMATICS AND COMPUTING↗

Role of electron correlation on the adenine dimer interaction for non-equilibrium geometries: a benchmark Quantum Monte Carlo study

The accurate description of non-covalent interactions is critical for understanding the structure, dynamics, and eventual function of biomolecules. The adenine dimer serves as a benchmark system for computational methods due to its role in nucleic acid structures and its rich conformational landscape. In this study, we employ benchmark diffusion quantum Monte Carlo (DMC) methods to investigate the relative energies and role of electron correlation on a set of adenine dimer conformations generated via a search of the potential energy landscape using the global optimizer algorithm. Relative DMC energies are compared against a wide range of density functional theory (DFT) approximation results. We find that although most of the DFT functionals perform well for low-energy structures, their accuracy varies significantly for higher-energy conformations, including stacked and T-shaped structures. A large fraction of the variation is due to the treatment of the van der Waals interaction. BLYP, B3LYP, and PBE0 significantly improve with added D4 dispersion, while the recent r2SCAN-D4 and ωB97M-V functionals show the least scatter and closest agreement with the DMC. These findings highlight the delicate nature of these interactions in biomolecular systems and provide guidance for simulations of their structure and dynamics and for the development of machine learned interatomic potentials.

Washburn, Laurel [ORNL] (ORCID:0000000324179335)↗

Matrix Product (GEMM) Performance Data from GPUs

Timing data for mixed precision GEMM matrix product operations on several GPU models, including NVIDIA V100 and A100, AMD MI100 and Intel P580. Also data from machine learning model training on this data using Scikit-learn.

97 MATHEMATICS AND COMPUTING↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Scientific machine learning benchmarks

Deep learning has transformed the use of machine learning technologies for the analysis of large experimental datasets. In science, such datasets are typically generated by large-scale experimental facilities, and machine learning focuses on the identification of patterns, trends and anomalies to extract meaningful scientific insights from the data. In upcoming experimental facilities, such as the Extreme Photonics Application Centre (EPAC) in the UK or the international Square Kilometre Array (SKA), the rate of data generation and the scale of data volumes will increasingly require the use of more automated data analysis. Furthermore, at present, identifying the most appropriate machine learning algorithm for the analysis of any given scientific dataset is a challenge due to the potential applicability of many different machine learning frameworks, computer architectures and machine learning models. Historically, for modelling and simulation on high-performance computing systems, these issues have been addressed through benchmarking computer applications, algorithms and architectures. Extending such a benchmarking approach and identifying metrics for the application of machine learning methods to open, curated scientific datasets is a new challenge for both scientists and computer scientists. Here, we introduce the concept of machine learning benchmarks for science and review existing approaches. As an example, we describe the SciMLBench suite of scientific machine learning benchmarks.

42 ENGINEERING↗

Accelerating the density-functional tight-binding method using graphical processing units

Acceleration of the density-functional tight-binding (DFTB) method on single and multiple graphical processing units (GPUs) was accomplished using the MAGMA linear algebra library. Herein two major computational bottlenecks of DFTB ground-state calculations were addressed in our implementation: the Hamiltonian matrix diagonalization and the density matrix construction. The code was implemented and benchmarked on two different computer systems: (1) the SUMMIT IBM Power9 supercomputer at the Oak Ridge National Laboratory Leadership Computing Facility with 1–6 NVIDIA Volta V100 GPUs per computer node and (2) an in-house Intel Xeon computer with 1–2 NVIDIA Tesla P100 GPUs. The performance and parallel scalability were measured for three molecular models of 1-, 2-, and 3-dimensional chemical systems, represented by carbon nanotubes, covalent organic frameworks, and water clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Integrated High-performance Computing and Digital Real-time Simulation Testbed to Benchmark Closed-loop Load Shedding Algorithms in Power Systems

An integrated testbed using digital real-time simulator (DRTS) and a high-performance computing (HPC) cluster is presented here to compare speed and performance of computational schemes to mitigate time-critical issues in electric power systems. The first approach in this testbed validation is taken by running a set of closed-loop load shedding algorithms to compare and contrast two paradigms of arresting cascading failure propagation. Two algorithms involve solving DC and AC power flow model-based optimization problems to compute load shedding at different buses, while a model-based stochastic search using parallel computing provides a viable alternative. The algorithms are implemented in the DRTS-HPC testbed for the IEEE 14-bus benchmark transmission system. As a proof of the concept, simulation results are presented for implementation of closed-loop load-shedding algorithms for cascading failures in the DRTS-HPC testbed

24 POWER TRANSMISSION AND DISTRIBUTION↗

Development of an Improved RELAP5-3D Model for the High Temperature Test Facility

High-temperature gas-cooled reactors (HTGRs) are rapidly approaching deployment. Confidence in transient analysis of these systems for design, optimization, and licensing calculations requires modeling and simulation tools that have been validated against data relevant to HTGR conditions. The High Temperature Test Facility (HTTF) is an integral effects thermal hydraulics test facility for prismatic HTGRs. In spring and summer of 2019, HTTF was used for a series of experiments that now serve as the basis for the OECD/NEA Thermal Hydraulic Code Validation Benchmark for High Temperature Gas-Cooled Reactors using HTTF Data (HTGR T/H Benchmark). This benchmark contains problems for systems code, computational fluid dynamics (CFD), and coupled systems code/CFD modeling representing lower plenum mixing and both the depressurized and pressurized conduction cooldown (DCC and PCC respectively) transients. Benchmark problems include exercises for code-to-code and code-to-data comparisons as well as an exercise for error scaling between HTTF and the Modular High Temperature Gas-Cooled Reactor, which serves as the basis for the HTTF design. Previous analysis as part of the HTGR T/H benchmark used a RELAP5-3D model developed at Idaho National Laboratory (INL) and demonstrated an ability to reproduce trends in the measured data but difficulties reproducing experimental values within their uncertainty. These difficulties were largely attributed to assumptions made during the development of the initial RELAP5-3D model, which predated the HTTF experiments. A significant cause of difficulty reproducing the measured temperatures may be the radial nodalization of the previous RELAP5-3D model. The new model provides a finer nodalization to assess the impact of radial nodalization and allows for asymmetric heating within the core, which was a feature of multiple HTTF experiments. In this paper, we present the new RELAP5-3D model of HTTF. In addition to describing the new model, this paper compares the new and old models and provides results for a full-power steady state, a DCC, and a PCC in HTTF. These analyses are based on the code-to-code comparison exercises for the DCC and PCC problems of the HTGR T/H benchmark. We present the results of these exercises from the new model and compare them to the results of the old model.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimal Realization of Yang–Baxter Gate on Quantum Computers

Quantum computers provide a promising method to study the dynamics of many-body systems beyond classical simulation. On the other hand, the analytical methods developed and results obtained from the integrable systems provide deep insights on the many-body system. Quantum simulation of the integrable system not only provides a valid benchmark for quantum computers but is also the first step in studying integrable-breaking systems. The building block for the simulation of an integrable system is the Yang–Baxter gate. It is vital to know how to optimally realize the Yang–Baxter gates on quantum computers. Based on the geometric picture of the Yang–Baxter gates, the optimal realizations of two types of Yang–Baxter gates with a minimal number of controlled NOT (CNOT) or gates are presented. It is also shown how to systematically realize the Yang–Baxter gates via the pulse control. The different realizations on IBM quantum computers are tested and compared. It is found that the pulse realizations of the Yang–Baxter gates always have a higher gate fidelity compared to the optimal CNOT or realizations. On the basis of the above optimal realizations, the simulation of the Yang–Baxter equation on quantum computers is demonstrated. Finally, these results provide a guideline and standard for further experimental studies based on the Yang–Baxter gate.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Random circuit block-encoded matrix and a proposal of quantum LINPACK benchmark

The LINPACK benchmark reports the performance of a computer for solving a system of linear equations with dense random matrices. Although this task was not designed with a real application directly in mind, the LINPACK benchmark has been used to define the list of TOP500 supercomputers since the debut of the list in 1993. We propose that a similar benchmark, called the quantum LINPACK benchmark, could be used to measure the whole machine performance of quantum computers. The success of the quantum LINPACK benchmark should be viewed as the minimal requirement for a quantum computer to perform a useful task of solving linear algebra problems, such as linear systems of equations. We propose an input model called the Random Circuit Block-Encoded Matrix (RACBEM), which is a proper generalization of a dense random matrix in the quantum setting. The RACBEM model is efficient to be implemented on a quantum computer and can be designed to optimally adapt to any given quantum architecture, with relying on a black-box quantum compiler. Besides solving linear systems, the RACBEM model can be used to perform a variety of linear algebra tasks relevant to many physical applications, such as computing spectral measures, time series generated by a Hamiltonian simulation, and thermal averages of the energy. We implement these linear algebra operations on IBM Q quantum devices as well as quantum virtual machines, and demonstrate their performance in solving scientific computing problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Immortal rays: Rethinking random ray neutron transport on GPU architectures

The Random Ray Method (TRRM) is a recently developed adaptation of the Method of Characteristics for neutral particle transport simulations. TRRM has demonstrated excellent performance on 3D nuclear reactor benchmark problems using CPU-based compute systems. When porting to GPU-based systems, however, new performance challenges arise that are unique to processors targeting massive fine-grained parallelism. For smaller problems, or for large problems that are domain decomposed across many computational nodes, the problem size per node has insufficient parallelism to saturate GPU node resources, thus greatly limiting speedup. In this study, we report on a newly developed “immortal ray” variant of TRRM. Here, the immortal ray technique exposes significantly more fine-grained parallelism by fundamentally reformulating the numerical details of ray discretization, resulting in performance tradeoffs with significant overall benefit on GPUs. For very small 2D simulation problems we found the new immortal ray variant allowed for up to a 4.4x speedup when run on a single GPU. For larger 3D simulation problems we found the new variant improved strong scaling by 3x when run on the Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Scalable Risk Assessment of Rare Events in Power Systems With Uncertain Wind Generation and Loads

Risk assessment of rare events has become increasingly important in power system planning and operation with the increasing integration of renewable energy and the presence of system uncertainties. However, quantifying the risk posed by rare events via the traditional method, i.e., Monte Carlo sampling (MCS), incurs substantial computational expense stemming from the vast ensemble of power flow simulations. To accelerate the assessment, this paper proposes a Deep Neural Network (DNN)-kernelized vector-valued Gaussian Process (VVGP) approach with excellent computational efficiency while maintaining high accuracy. Consequently, serving as a surrogate model for the power flow solver, the DNN-kernelized VVGP enables significantly faster but accurate risk assessment compared to the power flow solver. The developed surrogate model evaluates low-order N - k events that contain more than 90% instances by adeptly capturing the topological features while the high-order N - k events are assessed via a power flow solver, thereby striking a balance between computational efficiency and uncertainty quantification accuracy. Moreover, the model incorporates a Support Vector Machine (SVM) classifier to resample concerning low-probability tail events to counteract the biases potentially introduced during the DNN-kernelized VVGP evaluations. Simulations conducted on the modified IEEE 24-bus, 118-bus, and European 1354-bus systems demonstrate that the proposed method maintains the accuracy benchmark set by MCS while significantly reducing computational demands in large-scale power systems as compared to other state-of-the-art methods.

17 WIND ENERGY↗

SCALE 6.3 Validation: Radiation Shielding

Safe and reliable use of scientific and engineering computer codes requires validation for the types of applications in which they will be used. An example in the nuclear reactor engineering and licensing field is radiation transport employed in shielding analyses. The validity of computer codes for shielding applications is demonstrated in this report for SCALE version 6.3.0. Representative benchmarks corresponding to shielding analyses are selected for the validation study. Typical measurement results analyzed from these benchmarks include neutron fluxes, detector count rates, detector energy response functions, neutron and gamma dose rates, neutron activation rates and activities, neutron leakage fluxes, and skyshine dose rates. Thousands of points of comparison between measurement and calculation are presented in this work. Other than rare outliers typically explained by either a lack of information or large uncertainties in the experiment conditions, material, or dimensions, the Monaco with Automated Variance Reduction using Importance Calculations (MAVRIC) radiation transport computer code with built-in variance reduction methods distributed with the SCALE computer code system agrees well with the measurement results. In selected benchmarks, MAVRIC is also compared to Monte Carlo N- Particle® (MCNP® ) 1 calculations. Both computer codes generally agree well within the estimated uncertainties. With the release of SCALE 6.3.0, Shift was integrated as an alternative transport solver in MAVRIC, denoted MAVRIC-Shift. Although the traditional MAVRIC using Monaco was used primarily in this validation study, many results have also been generated using MAVRIC-Shift. Agreement between MAVRIC-Monaco and MAVRIC-Shift is generally very good. The benchmarks presented in this report were obtained from reliable sources such as the International Criticality Safety Benchmark Evaluation Project Handbook, the Shielding Integral Benchmark Archive & Database, and other shielding validation work found in the literature. Additional datapoints and benchmarks will be added to future versions of this report to expand the shielding validation suite.

61 RADIATION PROTECTION AND DOSIMETRY↗

OpenMxP-Opensource Mixed Precision Computing

This is an opensource library for benchmarking the system's GPU mixed precision capabilities. The software calculates solution of the system of linear equation in 64bit accuracy using mixed precision techniques and iterative refinement. Original benchmark designed is done by ICL, and it is name HPL-MxP (HPL-AI)

Lu, Hao↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

Practical Scalability of LuGo: Benchmarking the HHL Algorithm Using an Enhanced QPE Algorithm

The HHL algorithm is a prominent quantum algorithm that offers exponential speedup over its classical counterparts for solving a system of linear equations. However, synthesizing and executing HHL circuits demand significant computational resources from both classical and quantum systems. In this paper, we benchmark the HHL algorithm using the optimized Quantum Phase Estimation (QPE) generation algorithm, LuGo \cite{lu2025lugo}, to enhance its scalability and efficiency. We leverage the National Energy Research Scientific Computing Center's (NERSC) Perlmutter supercomputer to evaluate the scalability of generating HHL circuits and to measure the time to simulate the generated circuits. Additionally, we provide a comprehensive analysis of the algorithm's performance on various state-of-the-art superconducting and trapped-ion quantum devices, including studies on qubit connectivity, fidelity comparisons, and hardware compatibility and robustness. Our results offer preliminary insights into potential practical applications of the HHL algorithm enabled by LuGo and the performance of various types of quantum hardware.

Lu, Chao [ORNL] (ORCID:0000000179346933)↗