Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer Hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Tensorized Interior Radiative Heat Transfer for a Scalable and Calibrated Building Energy Simulator

Building energy simulation is a critical tool for developing and testing advanced control strategies, such as Reinforcement Learning (RL), to provide demand flexibility and affordable energy costs. The recently introduced Smart Buildings Control Suite (sbsim) provides a lightweight, scalable, and data-calibrated simulation environment based on a 2D finite-difference model. However, the initial model primarily focused on conductive and convective heat transfer, neglecting the significant impact of long-wave radiative heat exchange between interior surfaces. This paper presents a significant extension to the sbsim framework by incorporating a physically-grounded model for interior radiative heat transfer. Our primary contribution is the development and integration of a fully tensorized radiative heat transfer module, which preserves the computational efficiency and scalability of the original simulator. This was achieved by developing a pipeline for view factor calculation, including an algorithm to identify directly seeing surfaces within complex floor plans, and formulating the net radiation equations for efficient execution on modern hardware accelerators. We validate the numerical accuracy of our tensorized implementation by comparing its results against a traditional iterative approach, demonstrating identical outcomes. This enhancement increases the physical fidelity of sbsim, enabling more accurate training of RL agents for building energy optimization.

Ham, Sang woo↗

Beyond real: alternative unitary cluster Jastrow models for molecular electronic structure calculations on near-term quantum computers

Near-term quantum devices require wavefunction ansätze that are expressive while also of shallow circuit depth in order to both accurately and efficiently simulate molecular electronic structure. While the unitary coupled cluster ansatz (e.g., UCCSD) has become a standard, the high gate count associated with the implementation of this limits its feasibility on noisy intermediate-scale quantum (NISQ) hardware. k -Fold unitary cluster Jastrow (uCJ) ansätze mitigate this challenge by providing O( kN 2 ) circuit scaling and favorable linear depth circuit implementation. Previous work has focused on the real orbitalrotation (Re-uCJ) variant of uCJ, which allows an exact (Trotter-free) implementation. Here we extend and generalize the k -fold uCJ framework by introducing two new variants, Im-uCJ and g-uCJ, which incorporate imaginary and fully complex orbital rotation operators, respectively. Similar to Re-uCJ, both of the new variants achieve quadratic gate-count scaling. Our results focus on the simplest k = 1 model, and show that the uCJ models frequently maintain energy errors within chemical accuracy (∼1 kcal mol −1 ). Both g-uCJ and Im-uCJ are more expressive in terms of capturing electron correlation and are also more accurate than the earlier Re-uCJ ansatz. We further show that Im-uCJ and g-uCJ circuits can also be implemented exactly, without any Trotter decomposition. Numerical tests using k = 1 on H 2 , H 3 + , Be 2 , C 2 H 4 , C 2 H 6 and C 6 H 6 in various basis sets confirm the practical feasibility of these shallow Jastrow-based ansätze for applications on near-term quantum hardware.

Tkachenko, Nikolay V. [University of California, B↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

COSMIC DAWN: Distributed Analysis of Wireless at Nextscale

Distributed Analysis of Wireless at Nextscale (DAWN) is a novel simulation framework for large-scale design-space exploration (DSE) of unmodified software-defined radio (SDR) applications interacting in a scalable, high-fidelity, virtual physics environment. The software-defined nature of the coupled software-physics simulation leverages hardware emulation to permit in-depth examination and modification of not only the electromagnetic environment, including each signal in flight, but also the precise state of system software and components. DAWN supports modular, customizable physics environments allowing realistic propagation effects so that computationally efficient empirical models, reduced order/surrogate models, or large-scale, high-fidelity, site-specific simulations can be used as a propagation medium based on scenario requirements. This paper introduces DAWN’s design and initial implementation, detailing key architectural components, including the Physics Realization Engine (PhyRE), Runtime Infrastructure for Simulation Environments (RISE), and the design space exploration (DSE) suite. It concludes with demonstrations using unmodified 4G/LTE software available from srsRAN on computing resources ranging from a small cluster to ORNL’s Frontier Exascale system.

Wise, Mike [ORNL] (ORCID:0000000266120641)↗

Developing an Energy-Conscious Traffic Signal Control System for Optimized Fuel Consumption in Connected Vehicle Environments

The project titled “Developing an Energy-Conscious Traffic Signal Control System for Optimized Fuel Consumption in Connected Vehicle Environments” addresses energy-related challenges associated with adaptive traffic control systems by integrating connected vehicles (CV) and connected infrastructure (CI). The system developed in this project, a CV-based adaptive traffic control system, aims to improve fuel consumption in mixed traffic environments by capitalizing on emerging CV and CI communication technologies, as well as leveraging recent advances in Artificial Intelligence (AI), optimization, and edge computing. The system was tested at the MLK Smart Corridor, an urban testbed managed by the University of Tennessee at Chattanooga (UTC) and the City of Chattanooga. The system was validated through extensive simulations, both Software-in-the-Loop (SILS) and Hardware-in-the-Loop (HILS), and was further implemented and tested in real-world conditions at several intersections along the corridor. The Fuel Consumption Performance Index (FC-PI) and the Ecological Performance Index (Eco-PI) were developed as the key components for evaluating the system’s impact on fuel consumption and emissions. These metrics provided a comprehensive means of understanding the impact of traffic signal control optimization in mixed traffic environments. The report presents an in-depth analysis of the Eco-PI, FC-PI, adaptive traffic control system integration, and the testing and field implementation of the system. The results demonstrate significant reductions in fuel consumption and emissions, showcasing the system’s capability to contribute to more sustainable urban traffic management. The report also documents the challenges encountered and recommendations for scaling and further improving the system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

Trigonometric continuous-variable gates and hybrid quantum simulations of the sine-Gordon model

Hybrid qubit-qumode quantum computing platforms provide a natural setting for simulating interacting bosonic quantum field theories. However, existing continuous-variable gate constructions rely predominantly on polynomial functions of canonical quadratures. In this work, we introduce a complementary universality paradigm based on trigonometric continuous-variable gates, which enable a Fourier-like representation of bosonic operators and are particularly well suited for periodic and non-perturbative interactions. We present an ancilla-based framework for implementing trigonometric gates with arguments given by arbitrary Hermitian functions of qumode quadratures. The protocol yields unitary gates deterministically, and non-unitary gates through probabilistic post-selection. As a concrete application, we develop a hybrid qubit-qumode quantum simulation of the lattice sine-Gordon model. Using these gates, we prepare ground states via quantum imaginary-time evolution, simulate real-time dynamics, compute time-dependent vertex two-point correlation functions, and extract quantum kink profiles under topological boundary conditions. Our results demonstrate that trigonometric continuous-variable gates provide a physically natural framework for simulating interacting field theories on near-term hybrid quantum hardware, while establishing a parallel route to universality beyond polynomial gate constructions. We expect that the trigonometric gates introduced here to find broader applications, including quantum simulations of condensed matter systems, quantum chemistry, and biological models.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Phase-based velocity extraction method for photonic Doppler velocimetry with potential higher time resolution

We present an extension of the [Takeda et al., J. Opt. Soc. Am. 72, 156 (1982)] phase extraction method to heterodyne photonic Doppler velocimetry applications. The method yields results equivalent to those obtained by the short-time Fourier transform (STFT), while offering potential improvements in time resolution. Unlike STFT, which relies on window functions, such as the Hamming window, that emphasize central data points and diminish the influence of edges, the extended Takeda method utilizes all data uniformly. This uniform treatment allows for the derivation of empirical equations that directly relate velocity error to the actual time resolution rather than to the local analysis duration. The established equation provides a useful metric for both optimizing hardware configuration and guiding data analysis. Simulation and experimental results confirm that, for a given dataset, specifying a target time resolution yields consistent velocity errors for both methods. These findings underscore the Takeda method’s advantages, particularly its potential higher time resolution and reduced computational burden, making it a valuable tool for high-throughput applications such as laser dynamic compression experiments.

Computer simulation↗

Attention to quantum complexity

The imminent era of error-corrected quantum computing demands robust methods to characterize quantum state complexity from limited, noisy measurements. We introduce the Quantum Attention Network (QuAN), a classical artificial intelligence (AI) framework leveraging attention mechanisms tailored for learning quantum complexity. Inspired by large language models, QuAN treats measurement snapshots as tokens while respecting permutation invariance. Combined with our parameter-efficient miniset self-attention block, this enables QuAN to access high-order moments of bit-string distributions and preferentially attend to less noisy snapshots. We test QuAN across three quantum simulation settings: driven hard-core Bose-Hubbard model, random quantum circuits, and toric code under coherent and incoherent noise. QuAN directly learns entanglement and state complexity growth from experimental computational basis measurements, including complexity growth in random circuits from noisy data. In regimes inaccessible to existing theory, QuAN unveils the complete phase diagram for noisy toric code data as a function of both noise types, highlighting AI’s transformative potential for assisting quantum hardware.

Kim, Hyejin [Cornell Univ., Ithaca, NY (United Sta↗

Realistic Cost to Execute Practical Quantum Circuits using Direct Clifford+T Lattice Surgery Compilation

We report a resource estimation pipeline that explicitly compiles quantum circuits expressed using the Clifford+T gate set into a surface code lattice surgery instruction set. The cadence of magic state requests from the compiled circuit enables the optimization of magic state distillation and storage requirements in a post-hoc analysis. To compile logical circuits into lattice surgery operations, we build upon the open-source Lattice Surgery Compiler. The revised compiler operates in two stages: the first translates logical gates into an abstract, layout-independent instruction set; the second compiles these into local lattice surgery instructions that are allocated to hardware tiles according to a specified resource layout. The second stage retains logical parallelism while avoiding resource contention in the fault-tolerant layer, aiding realism. Additionally, users can specify dedicated tiles at which magic states are replenished, enabling resource costs from the logical computation to be considered independently from magic state distillation and storage. We demonstrate the applicability of our pipeline to large practical quantum circuits by providing resource estimates for the ground state estimation of molecules. Finally, we find that variable magic state consumption rates in real circuits can cause the resource costs of magic state storage to dominate unless production is varied to suit.

97 MATHEMATICS AND COMPUTING↗

ML–Enabled FPGA Framework for Fast Quantum State Discrimination in Mid-Circuit Measurement Regimes

Accurate and low-latency quantum state discrimination is essential for protocols involving mid-circuit measurement (MCM) and conditional feed-forward. In superconducting quantum systems, conventional readout pipelines transfer measurement data to host processors for post-processing, introducing millisecond-scale delays that far exceed qubit coherence times. To overcome this bottleneck, we present an in-situ machine learning (ML) inference engine implemented on an FPGA for real-time quantum state discrimination. Our design performs inference directly on digitized readout signals with 40 ns latency, supports both qubit and qutrit readout, and enables conditional operations without host-side intervention. This capability is critical for MCM and for feedback-driven protocols such as quantum error correction. We validate the system on superconducting transmon hardware, demonstrating robust discrimination fidelity across multiple qubit and qutrit channels. We further demonstrate conditional qutrit logic driven by FPGA-resident classification, highlighting the potential of low-latency ML-on-FPGA control for NISQ applications and scalable fault-tolerant quantum computing.

Vora, Neel [Lawrence Berkeley National Laboratory ↗

ArborX 2.0

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, ArborX uses linear BVH for its low construction cost and sufficient quality. ArborX implements both spatial and nearest-neighbor traversal algorithms. ArborX also provides several clustering algorithms (minimum spanning tree, DBSCAN, HDBSCAN*), interpolation using minimum least squares and ray tracing. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase.

Prokopenko, Andrey [Oak Ridge National Laboratory ↗

Error mitigation, optimization, and extrapolation on a trapped-ion testbed

Current noisy intermediate-scale quantum (NISQ) trapped-ion devices are subject to errors which can significantly impact the accuracy of calculations if left unchecked. A form of error mitigation called zero noise extrapolation (ZNE) can decrease an algorithm’s sensitivity to these errors without increasing the number of required qubits. Here we explore different methods for integrating this error mitigation technique into the Variational Quantum Eigensolver (VQE) algorithm for calculating the ground state of the HeH + molecule at 0.8 Å in the presence of experimental noise. Using the Quantum Scientific Computing Open User Testbed (QSCOUT) trapped-ion device, we test three methods of scaling noise for extrapolation: time stretching the two-qubit gates, scaling the sideband detuning parameter, and inserting two-qubit gate identity operations into the ansatz circuit. We find that time stretching and sideband detuning scaling fail to scale the noise on our particular hardware in a way that can be extrapolated to zero noise. Scaling our noise with global gate identity insertions and extrapolating after variational optimization, we achieve error suppression of 96.8%, resulting in an energy estimate within –0.004 ± 0.04 hartree of the ground state energy. This is an improvement, but still outside the chemical accuracy threshold of 0.0016 hartree. Furthermore, our results show that the efficacy of this error mitigation technique depends on choosing the correct implementation for a given device architecture.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Multi-purpose quantum laboratories from superconducting circuits

Superconducting circuits (SCs) are the cornerstone of modern quantum technology, enabling scalable computing through coherent control of macroscopic quantum states. Through a legacy that predates modern quantum computing, SCs have emerged as high-precision instruments for discovery. In this review, we highlight the role of SCs as general-purpose quantum laboratories, outlining the emerging landscape of correlated matter-circuit science. We review and unify the capabilities of superconducting quantum hardware across condensed matter, high energy and quantum information sciences. We trace the technical evolution of these architectures, illustrating how their foundational development has culminated in a toolkit for resolving the complexities of macroscopic quantum states.

Arora, Arpit [UCLA, Los Angeles (main); UCLA; Haim↗

S&TR September 2025: Computing Grand Challenge Turns 20

Livermore’s Computing Grand Challenge Program enters its 20th year with more unclassified high-performance computing (HPC) power than ever before. This unique, peer-reviewed competition awards HPC allocations on top supercomputers to multidisciplinary teams with high-impact projects. The Grand Challenge encourages researchers to innovate, pushes scientific discovery to new heights, improves the Laboratory’s HPC capabilities, and extends HPC accessibility to collaborators. Awardees must adapt to successive generations of HPC hardware and learn to run simulations at scale. The feature article spotlights three Grand Challenge teams whose research broke new ground in key scientific pursuits—the essence of dark matter, explosion-generated seismic waves, and protein interactions linked to cancer—while underscoring the importance of academic partnerships and considering the program’s future.

07 ISOTOPE AND RADIATION SOURCES↗

Understanding and Mitigating Coherence and Frequency Fluctuations in Superconducting Transmon Qubits

Transmon qubits are a cornerstone of superconducting quantum computing platforms. However‚ their frequency and coherence properties exhibit temporal fluctuations‚ leading to performance degradation in quantum processors over time. A common mitigation approach involves frequent recalibration‚ which‚ while effective‚ results in increased system downtime. Enhancing the long-term stability of transmon qubits is therefore critical for scalable and reliable quantum computing. In this study‚ we develop novel techniques for understanding the underlying mechanisms driving frequency and coherence fluctuations in fixed-frequency transmon qubits. We further explore strategies to mitigate these instabilities‚ aiming to improve overall system robustness. Our findings provide insights into optimizing superconducting quantum hardware for practical applications.

Roy, Tanay [Fermilab]↗