Search NASA⌕ Search

SEARCH · Search NASA

Results for “Hardware acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Keeping LAMMPS cutting edge

Since its inception 30 years ago, LAMMPS has grown to be a world-class molecular dynamics code and a cornerstone of computational materials science research. This project aimed to keep LAMMPS at the forefront of molecular dynamics simulations by adapting LAMMPS to the latest developments in machine learning technology and hardware. Initially, the project set out to provide a unified implementation of active learning for efficient training data generation in LAMMPS, but the research trajectory pivoted to address more immediate and impactful opportunities. On the hardware side, recent record-breaking molecular dynamics simulations were developed on the Cerebras wafer-scale AI chip, and this project has developed an interface between LAMMPS and the hardware-specific molecular dynamics code to accelerate and simplify development and user adoption. On the software side, PyTorch’s Ahead-of-Time (AOT) compilation features promised increased performance for state-of-the-art equivariant neural network potentials, and this project laid the groundwork for their adoption in LAMMPS, resulting in a nearly 20x acceleration in extreme cases. Combined with a comprehensive benchmark study of LAMMPS across all current exascale systems, this project has reinforced LAMMPS’s role as a versatile, high-performance tool for current and future materials science applications.

36 MATERIALS SCIENCE↗

Surrogate optimization of variational quantum circuits

Variational quantum eigensolvers are touted as a near-term algorithm capable of impacting many applications. However, the potential has not yet been realized, with few claims of quantum advantage and high resource estimates, especially due to the need for optimization in the presence of noise. Finding algorithms and methods to improve convergence is important to accelerate the capabilities of near-term hardware for VQE or more broad applications of hybrid methods in which optimization is required. To this goal, we look to use modern approaches developed in circuit simulations and stochastic classical optimization, which can be combined to form a surrogate optimization approach to quantum circuits. Using an approximate (classical CPU/GPU) state vector simulator as a surrogate model, we efficiently calculate an approximate Hessian, passed as an input for a quantum processing unit or exact circuit simulator. This method will lend itself well to parallelization across quantum processing units. We demonstrate the capabilities of such an approach with and without sampling noise and a proof-of-principle demonstration on a quantum processing unit utilizing 40 qubits.

Gustafson, Erik J. [RIACS, Mtn. View] (ORCID:00000↗

Beam Synchronous for the Rest of Us!

Fermilab’s Tevatron Clock (TCLK) infrastructure has been an integral part of the accelerator control network since the 1980’s. This 10MHz Manchester encoded protocol has enabled flexible, real-time event distribution for thousands of devices connected to the timing network with a high degree of reliability. Forthcoming upgrades to the Fermilab complex (PIP-II, LBNF, ACORN) necessitate higher levels of precision to maintain inter-bunch timing for Instrumentation and Control purposes. This presents as an opportunity to refine the event distribution protocol for tighter synchronization between machines, experiments, and eventually far-site operations. This paper outlines a method by which beam-synchronous events may be distributed through asynchronous serial protocols via integration with local LLRF and global PPS reference signals. This method is ideal for synchrotron machines with aggressive frequency sweeps (such as Fermilab's 38~53MHz Booster) and allows for precision timing to be maintained across machines without specialized hardware.

43 PARTICLE ACCELERATORS↗

High-$Q_0$ treatment of CEBAF 1.5 GHz SRF cavities

The Continuous Electron Beam Accelerator Facility (CEBAF) was the first large-scale accelerator to employ superconducting radiofrequency (SRF) cavities for continuous-wave operation. Ongoing research and development efforts continue to focus on increasing the intrinsic quality factor (Q 0 ) of these cavities in order to reduce cryogenic losses while maintaining operational gradients. Here, in this work, we report on the application of high-Q 0 surface treatments to single-cell and multicell C100 and C75 style 1.5 GHz niobium cavities used in the CEBAF accelerator. Nitrogen infusion and oxygen alloying via medium-temperature baking were applied under heat-treatment constraints relevant to existing cavity hardware. Both processes yielded substantial improvements in Q 0 at moderate accelerating gradients, achieving values of approximately 2 × 10 10 at 2.07 K and 20 MV/m. The effectiveness of nitrogen infusion at reduced annealing temperatures and the successful extension of oxygen alloying to multicell cavities are demonstrated. These results establish viable pathways for implementing high-Q 0 treatments in CEBAF-compatible cavities and support future integration into cryomodules for reduced operational cryogenic load.

Dhakal, Pashupati [Thomas Jefferson National Accel↗

Application of performance portability solutions for GPUs and many-core CPUs to track reconstruction kernels

Next generation High-Energy Physics (HEP) experiments are presented with significant computational challenges, both in terms of data volume and processing power. Using compute accelerators, such as GPUs, is one of the promising ways to provide the necessary computational power to meet the challenge. The current programming models for compute accelerators often involve using architecture-specific programming languages promoted by the hardware vendors and hence limit the set of platforms that the code can run on. Developing software with platform restrictions is especially unfeasible for HEP communities as it takes significant effort to convert typical HEP algorithms into ones that are efficient for compute accelerators. Multiple performance portability solutions have recently emerged and provide an alternative path for using compute accelerators, which allow the code to be executed on hardware from different vendors. We apply several portability solutions, such as Kokkos, SYCL, C++17 std::execution::par, Alpaka, and OpenMP/OpenACC, on two mini-apps extracted from the mkFit project: p2z and p2r. These apps include basic kernels for a Kalman filter track fit, such as propagation and update of track parameters, for detectors at a fixed z or fixed r position, respectively. The two mini-apps explore different memory layout formats. We report on the development experience with different portability solutions, as well as their performance on GPUs and many-core CPUs, measured as the throughput of the kernels from different GPU and CPU vendors such as NVIDIA, AMD and Intel.

Kwok, Ka Hei Martin↗

Deployment and validation of predictive 6-dimensional beam diagnostics through generative reconstruction with standard accelerator elements

Understanding the 6-dimensional phase space distribution of particle beams is essential for optimizing accelerator performance. Conventional diagnostics such as use of transverse deflecting cavities offer detailed characterization but require dedicated hardware and space. Generative phase space reconstruction (GPSR) methods have shown promise in beam diagnostics, yet prior implementations still rely on such components. Here we present the first experimental implementation and validation of the GPSR methodology, realized by the use of standard accelerator elements including accelerating cavities and dipole magnets, to achieve complete 6-dimensional phase space reconstruction. Through simulations and experiments at the Pohang Accelerator Laboratory X-ray Free Electron Laser facility, we successfully reconstruct complex, nonlinear beam structures. Furthermore, we validate the methodology by predicting independent downstream measurements excluded from training, revealing the reconstruction closely resembling ground truth. This advancement establishes a pathway for predictive diagnostics across beamline segments while reducing hardware requirements and expanding applicability to various accelerator facilities.

Kim, Seongyeol [Pohang Univ. of Science and Techno↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

DDCP framework

DDCP protocol software 1.0 This repository contains the C++ implementation of version 1.x of the Distributed Data Communications Protocol (DDCP). DDCP provides request/reply, feature discovery, data transfer, control, interrupt, and transaction support for communicating with accelerator instrumentation over UDP. The standard server port is 65000. The framework is a source dependency for services that communicate directly with DDCP hardware. It is not a deployable service by itself.

Joshi, Shreya [Fermi National Accelerator Laborato↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

Relief Zones Enhance the Durability of Ultrathin Membranes in Electrochemical Conversion Devices

Premature failures in electrochemical conversion systems often result when membrane electrode assemblies (MEAs) use ultrathin (≤15 μm-thick) polymer electrolyte membranes, susceptible to mechanical degradation from stress concentrations arising from device-level integration. Herein, relief zones were developed to mitigate mechanical degradation by alleviating excess and nonuniform compression across active areas. Relief zones, created through ablation of carbonaceous diffusion media, enable seamless adaptation across MEA dimensions without need for hardware modifications. Demonstrated using fuel cells as a case study, accelerated stress tests revealed a 6-fold lifetime improvement (∼1500 h) compared to conventional edge-protected MEAs, decoupling device-level engineering effects from material limitations.

accelerated stress test↗

Data-driven analysis to understand GPU hardware resource usage of optimizations

With heterogeneous systems, the number of GPUs per chip increases to provide computational capabilities for solving science at a nanoscopic scale. However, low utilization for single GPUs defies the need to invest more money in expensive accelerators. Although related work develops optimizations to improve application performance, none studies how these optimizations impact hardware resource usage or average GPU utilization. Here, this paper takes a data-driven analysis approach in addressing this gap by (1) characterizing how hardware resource usage affects device utilization, execution time, or both, (2) presenting a multiobjective metric to identify important application-device interactions that can be optimized to improve device utilization and application performance jointly, (3) studying hardware resource usage behaviors of several optimizations for a benchmark application, and finally (4) identifying optimization opportunities for several scientific proxy applications based on their hardware resource usage behaviors. Furthermore, we demonstrate the applicability of our methodology by applying the identified optimizations to a proxy application, which improves the execution time, device utilization, and power consumption by up to 29.6%, 5.3% and 26.5% respectively.

Computer science↗

Calibration and Measurement techniques in the LLRF systems of the Fermilab PIP-II Linac

he PIP-II Accelerator is an 800 MeV superconducting Linac in the injection chain of the Fermilab accelerator complex. The LLRF systems are a based on two different hardware platforms controlling a variety of cavity types and resonance control systems including temperature, pneumatic and piezzo tuners. The various calibrations required prior to beam operation include, signal power, gradient, amplifier characterization, cavity Q measurement and detune constants. Measurements such as piezzo capacitance, cavity piezo transfer function help in determining tuner health and in devising microphonics control strategies. These measurement and calibration methods of the PIP-II LLRF system are discussed here.

Varghese, P.↗

Integrating quantum computing resources into scientific HPC ecosystems

Quantum Computing (QC) offers significant potential to enhance scientific discovery in fields such as quantum chemistry, optimization, and artificial intelligence. Yet QC faces challenges due to the noisy intermediate-scale quantum era’s inherent external noise issues. Here, this paper discusses the integration of QC as a computational accelerator within classical scientific high-performance computing (HPC) systems. By leveraging a broad spectrum of simulators and hardware technologies, we propose a hardware-agnostic framework for augmenting classical HPC with QC capabilities. Drawing on the HPC expertise of the Oak Ridge National Laboratory (ORNL) and the HPC lifecycle management of the Department of Energy (DOE), our approach focuses on the strategic incorporation of QC capabilities and acceleration into existing scientific HPC workflows. This includes detailed analyses, benchmarks, and code optimization driven by the needs of the DOE and ORNL missions. Our comprehensive framework integrates hardware, software, workflows, and user interfaces to foster a synergistic environment for quantum and classical computing research. This paper outlines plans to unlock new computational possibilities, driving forward scientific inquiry and innovation in a wide array of research domains.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Quantum Circuit Partitioning for Scalable Noise-Aware Quantum Circuit Re-Synthesis

Re-synthesis techniques are utilized to optimize the quantum circuit. To enable scalable re-synthesis a divide-and-conquer approach is adopted that partitions the circuit into smaller blocks, which are optimized independently. Several algorithms have been proposed to minimize the block number while maximizing the gate count of each block. However, they vary in their performance and may not yield the highest output fidelity. We propose a reinforcement learning-based quantum circuit partitioning framework that incorporates the physical properties of the quantum hardware to maximize the output fidelity post-quantum circuit optimization. To accelerate the training, we also propose a noise injection method that enables on-the-fly optimization in the reinforcement learning environment, independent of the adopted optimization/re-synthesis method at the block level. We evaluate our approach compared to different partitioning techniques using various quantum benchmarks executed on IBM Q Hanoi quantum computer.

Charrwi, Mohammad Walid↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-power test of a C-band linear accelerating structure with an RFSoC-based LLRF system

Normal conducting linear particle accelerators consist of multiple rf stations with accelerating structure cavities. Low-level rf (LLRF) systems are employed to set the phase and amplitude of the field in the accelerating structure and to compensate for the pulse-to-pulse fluctuation of the rf field in the accelerating structures with a feedback loop. The LLRF systems are typically implemented with analog rf mixers, heterodyne-based architectures, and discrete data converters. There are multiple rf signals from each of the rf stations, so the number of rf channels required increases rapidly with multiple rf stations. With a large number of rf channels, the footprint, component cost, and system complexity of the LLRF hardware will increase significantly. To meet the design goals of being compact and affordable for future accelerators, we have designed the next-generation LLRF (NG-LLRF) with a higher integration level based on RFSoC technology. The NG-LLRF system samples rf signals directly and performs rf mixing digitally. Further, the NG-LLRF has been characterized in loopback mode to evaluate the performance of the system and has also been tested with a standing-wave accelerating structure, a prototype for the Cool Copper Collider (C 3 ) with a peak rf power level up to 16.45 MW. The loopback test demonstrated amplitude fluctuation below 0.15% and phase fluctuation below 0.15°, which are considerably better than the requirements of C 3 . The rf signals from the different stages of the accelerating structure at different power levels are measured by the NG-LLRF, which will be critical references for the control algorithm designs. The NG-LLRF also offers flexibility in waveform modulation, so we have used rf pulses with various modulation schemes, which could be useful for controlling some of the rf stations in accelerators. In this paper, the high-power test results at different stages of the test setup will be summarized, analyzed, and discussed.

47 OTHER INSTRUMENTATION↗