Search NASA⌕ Search

SEARCH · Search NASA

Results for “GPU Computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

hls4ml

hls4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where hls4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

NeuroSEM: A hybrid framework for simulating multiphysics problems by coupling PINNs and spectral elements

Multiphysics problems that are characterized by complex interactions among fluid dynamics, heat transfer, structural mechanics, and electromagnetics, are inherently challenging due to their coupled nature. While experimental data on certain state variables may be available, integrating these data with numerical solvers remains a significant challenge. Physics-informed neural networks (PINNs) have shown promising results in various engineering disciplines, particularly in handling noisy data and solving inverse problems in partial differential equations (PDEs). However, their effectiveness in forecasting nonlinear phenomena in multiphysics regimes, particularly involving turbulence, is yet to be fully established. Here, this study introduces NeuroSEM, a hybrid framework integrating PINNs with the highfidelity Spectral Element Method (SEM) solver, Nektar++. NeuroSEM leverages the strengths of both PINNs and SEM, providing robust solutions for multiphysics problems. PINNs are trained to assimilate data and model physical phenomena in specific subdomains, which are then integrated into the Nektar++ solver. We demonstrate the efficiency and accuracy of NeuroSEM for thermal convection in cavity flow and flow past a cylinder. The framework effectively handles data assimilation by addressing those subdomains and state variables where the data is available. We applied NeuroSEM to the Rayleigh-B´enard convection system, including cases with missing thermal boundary conditions and noisy datasets. Finally, we applied the proposed NeuroSEM framework to real particle image velocimetry (PIV) data to capture flow patterns characterized by horseshoe vortical structures. Our results indicate that NeuroSEM accurately models the physical phenomena and assimilates the data within the specified subdomains. The framework’s plug-and-play nature facilitates its extension to other multiphysics or multiscale problems. Furthermore, NeuroSEM is optimized for efficient execution on emerging integrated GPU-CPU architectures. This hybrid approach enhances the accuracy and efficiency of simulations, making it a powerful tool for tackling complex engineering challenges in various scientific domains.

42 ENGINEERING↗

To Exascale and Beyond—The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM), a Performance Portable Global Atmosphere Model for Cloud-Resolving Scales

The new generation of heterogeneous CPU/GPU computer systems offer much greater computational performance but are not yet widely used for climate modeling. One reason for this is that traditional climate models were written before GPUs were available and would require an extensive overhaul to run on these new machines. In addition, even conventional “high–resolution” simulations don't currently provide enough parallel work to keep GPUs busy, so the benefits of such overhaul would be limited for the types of simulations climate scientists are accustomed to. The vision of the Simple Cloud-Resolving Energy Exascale Earth System (E3SM) Atmosphere Model (SCREAM) project is to create a global atmospheric model with the architecture to efficiently use GPUs and horizontal resolution sufficient to fully take advantage of GPU parallelism. After 5 years of model development, SCREAM is finally ready for use. In this paper, we describe the design of this new code, its performance on both CPU and heterogeneous machines, and its ability to simulate real-world climate via a set of four 40 day simulations covering all 4 seasons of the year.

54 ENVIRONMENTAL SCIENCES↗

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE↗

Direct sensitivity analysis on the parameterization of crystal plasticity models

Various methods for calibrating crystal plasticity finite element (CPFE) models lead to non-unique input parameter values, which subsequently introduce uncertainty in the predicted mechanical response. Sensitivity analysis (SA) conducted on crystal plasticity models is used to identify how variability in these parameters contribute to output uncertainty. Traditional SA on CPFE parameters uses simplified surrogate models to save computational time. However, the accuracy of the surrogate models depends on the quantity of training data used, and any modeling error can propagate into the SA results, potentially affecting their reliability. In this work, the elementary effects test (EET) method, a global SA technique using direct CPFE simulations was employed, and the results obtained were compared with the First Order Second Moment (FOSM) method. ExaConstit, an open-source GPU-enabled CPFE code, was used to perform the simulations and direct SA. The EET method was accurately able to capture the non-linear effects of all the input parameters on the output and is a valuable approach for reliably attributing parameter sensitivities in CPFE models. Based on the results, efficient strategies to perform future parameter calibration and SA are discussed. Additionally, the SA trends observed in different single crystal orientations closely mirrored the activity of the slip systems.

Elementary Effects Test↗

Scaling deep learning for material imaging with a pseudo 3D model for domain transfer

The recent introduction of deep learning methods for image processing has greatly advanced the characterization of materials using three-dimensional (3D) X-ray imaging techniques. However, deep learning models often have difficulty performing consistently across images owing to unavoidable variations in imaging conditions, which create inconsistencies even for the same material. As a result, networks must frequently be retrained for new datasets, limiting their applicability and generalization. Thus, it is critical to reduce the variations between images to enable a single model to process multiple datasets. Herein, we introduce P3T-Net, a pseudo-3D domain transfer network that transfers diverse 3D images into a uniform domain before processing using deep learning models. Remarkably, P3T-Net enables the reuse of previously trained networks for processing new images and considerably reduces the computational cost of transferring 3D images across domains. These unique capabilities were demonstrated in the following scenarios: (i) image enhancement of fast scans for geological rock and hydrogen fuel cells, (ii) enhancement of images to match the quality of multi-source imaging for lithium-ion batteries, (iii) accurate segmentation of images captured under different conditions, and (iv) tera-scale 3D transfer (10 11 voxels) on a single GPU. Overall, the proposed approach addresses cross-domain inconsistencies across various materials and conditions, thereby enabling more robust and generalizable deep learning solutions for a wide range of material imaging tasks.

25 ENERGY STORAGE↗

If We Build Them, They Will Run: Automated HPC Apps Deployment and Profiling with eBPF in Cloud

The high performance computing (HPC) community is in a period of transition. The rise of AI/ML coupled with a changing landscape of resources deems portability a new metric of performance, and methods to move between on-premises and cloud environments and assess compatibility are paramount. Here we design and test a strategy for bridging the gap between traditional HPC and Kubernetes environments – first containerizing applications, providing automated orchestration to run studies, and packaging the setup with automated means to assess performance using low overhead eXtended Berkeley Packet Filter (eBPF) programs. We first assess different designs for eBPF collection, demonstrating a tradeoff between number of programs deployed on a node and overhead added. We develop 5 low overhead eBPF programs that combine with streaming ML models to assess CPU, futex, TCP, shared memory, and file access across four different builds of an HPC application for CPU and GPU. We use eBPF data to generate insights into the possible underlying etiology of scaling issues. We then assess compatibility of a well-known benchmark, HPCG, across matrices of micro-architectures and optimization levels (217 containers across 24 instance types and over 7500 runs). We provide to the community 30 applications to deploy in our automated setup and perform a scaling study from 4 to a maximum of 256 nodes for both CPU and GPU applications. Finally, we use our gained knowledge about performance to generate compatibility artifacts that are used by a newly developed Kubernetes controller to intelligently select instance type based on optimizing a figure of merit. Along with insights to scaling in this environment with a collection of applications and templates to work from, we provide an overall strategy for approaching HPC application deployment and image selection based on compatibility in cloud.

Computer science↗

GPU-Accelerated Analytic Simulation of Sparse Ionization Signal Formation in Pixelated Projection Detector

This paper presents a GPU-accelerated simulation package, TRED, for next-generation neutrino detectors with pixelated charge readout, leveraging community-driven software ecosystems to ensure adaptability and extensibility. We introduce two generic contributions: (i) an effective-charge representation based on Gaussian quadrature rules, in which the linear- interpolation factors for the field response inside each voxel are absorbed into the effective charge, and (ii) a sparse, block- binned tensor representation that enables efficient FFT-based computation of induced signals on readout electrodes for sparsely activated detector volumes. The former captures structure inside a voxel without dense sampling, while the latter achieves low memory usage and scalable runtime, as demonstrated in bench- mark studies. The underlying data representation is applicable to large-scale detectors and to other computational problems involving sparse activity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ristra Project FY23 L2 Milestone Report, Rev.1: MRT #8541: Multiphysics Scaling on EAS-3

The findings of this report were used to close out the ATDM milestone MRT# 8541, which was designed to demonstrate readiness of ATDM multiphysics codes for mission-relevant work on ATS-4, El Capitan. To this end, the closure criteria were to run a 3D shaped charge problem at scale up to 50% of the El Capitan early-access system, RZVernal (AMD Trento CPUs and AMD MI-250X GPUs), demonstrate scalability, and document challenges with the software stack and environment. LANL’s approach to this milestone was to test our modular software capability by developing an entirely new code, Moya, built upon our FleCSI framework. The physics capability and the GPU infrastructure needed for the shaped charge problem on GPUs was added to Moya, and the required calculations were performed at scale. Moya showed good scaling without any fine-tuning of GPU kernels; there is still significant room for performance enhancements, especially for the Legion backend. Tied up in this L2 milestone was a closeout of KPP-3s for the ECP ST Projects at LANL; this material will be covered in a separate document.

97 MATHEMATICS AND COMPUTING↗

Molecular Dynamics Simulation of Complex Reactivity with the Rapid Approach for Proton Transport and Other Reactions (RAPTOR) Software Package

Simulating chemically reactive phenomena such as proton transport on nanosecond to microsecond and beyond time scales is a challenging task. Ab initio methods are unable to currently access these time scales routinely, and traditional molecular dynamics methods feature fixed bonding arrangements that cannot account for changes in the system’s bonding topology. The Multiscale Reactive Molecular Dynamics (MS-RMD) method, as implemented in the Rapid Approach for Proton Transport and Other Reactions (RAPTOR) software package for the LAMMPS molecular dynamics code, offers a method to routinely sample longer time scale reactive simulation data with statistical precision. RAPTOR may also be interfaced with enhanced sampling methods to drive simulations toward the analysis of reactive rare events, and a number of collective variables (CVs) have been developed to facilitate this. Key advances to this methodology, including GPU acceleration efforts and novel CVs to model water wire formation are reviewed, along with recent applications of the method which demonstrate its versatility and robustness.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncovering grain and subgrain microstructure at the scale of additive manufacturing melt tracks with a scalable cellular automaton solidification model

Metal additive manufacturing, characterized by rapid solidification, yields refined grains with a distinctive cellular subgrain microstructure that plays a pivotal role in determining material properties. Due to the significant computational expense demanded to simulate the required physics with submicron spatial resolution, their numerical simulations have been limited to proof-of-concept studies to either 2D or small subregions of a melt pool. In this study, an open-source, scalable, solidification code, muMatScale, based on the cellular automaton method, has been developed to predict the grain and the underlying subgrain microstructure over an entire melt pool. The model incorporates flexible parallelization schemes, utilizing MPI and OpenMP GPU Offloading, in addition to appropriate multi-physics specific to non-equilibrium rapid solidification in AM. The impact of nucleation parameters on grain microstructures was investigated with a focus on grain size variations and morphology transitions. With selected nucleation parameters, the simulation predicted the grain size, subgrain morphology, crystallographic orientation, and microsegregation aligned with experimental measurements. The model demonstrates that epitaxial grain growth is a dominant factor at the melt pool boundary, influencing grain size variation under different grain sizes in the build plate while maintaining consistent primary dendrite arm spacing under identical thermal conditions. Here, the highly efficient numerical model enables large-scale simulations with a spatial resolution of 100 nm or less, unveiling unprecedented insights into thermal and solutal diffusion driven grain growth, and the subgrains with microsegregation within grains in 3D across scales. muMatScale will enable the linking of submicron length-scale microstructure to part-level material behavior by investigating fundamental solidification problems at the intercellular scale in many-track and many-layer builds.

36 MATERIALS SCIENCE↗

Multitarget Rydberg gates via spatial blockade engineering

Multi-target gates offer the potential to reduce gate depth in syndrome extraction for quantum error correction. Although neutral-atom quantum computers have demonstrated native multi-qubit gates, existing approaches that avoid additional control or multiple atomic species have been limited to single-target gates. We propose single-control-multi-target CZ^n gates on a single-species neutral-atom platform that require no extra control and have gate durations comparable to standard CZ gates. Our approach leverages tailored interatomic distances to create an asymmetric blockade between the control and target atoms. Using a GPU-accelerated pulse synthesis protocol, we design smooth control pulses for CZZ and CZZZ gates, achieving fidelities of up to 99.55% and $99.24\%$, respectively, even in the presence of simulated atom placement errors and Rydberg-state decay. Our approach is most effective for N=2 (CZZ) and N=3 targets (CZZZ); for larger N, increasing spatial crowding of the targets introduces significant challenges for maintaining the required blockade asymmetry. This work presents a practical path to implementing low-overhead multi-target gates in single-species neutral-atom systems, significantly reducing the resource overhead for syndrome extraction. To motivate the impact of these gates, we apply a greedy scheduling algorithm and we demonstrate that our proposed gates can reduce the number of atom reconfiguration costs by up to 50% for color code syndrome extraction of code distances greater than 5.

Stein, Samuel A.↗

Singularity-EOS: Performance Portable Equations of State and Mixed Cell Closures

We present Singularity-EOS, a new performance-portable library for equations of state and related capabilities. Singularity-EOS provides a large set of analytic equations of state, such as the Gruneisen equation of state, and tabulated equation of state data under a unified interface. It also provides support capabilities around these equations of state, such as Python wrappers, solvers for finding pressure-temperature equilibrium between multiple equations of state, and a unique modifier framework, allowing the user to transform a base equation of state, for example by shifting or scaling the specific internal energy. All capabilities are performance portable, meaning they compile and run on both CPU and GPU for a wide variety of architectures.

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Sensitivity of Simulations of Double-detonation Type Ia Supernovae to Integration Methodology

Abstract We study the coupling of hydrodynamics and reactions in simulations of the double-detonation model for Type Ia supernovae. When assessing the convergence of simulations, the focus is usually on spatial resolution; however, the method of coupling the physics together as well as the tolerances used in integrating a reaction network also play an important role. In this paper, we explore how the choices made in both coupling and integrating the reaction portion of a simulation (operator/Strang splitting versus the simplified spectral deferred corrections method we introduced previously) influences the accuracy, efficiency, and nucleosynthesis of simulations of double detonations. We find no need to limit reaction rates or reduce the simulation time step to the reaction timescale. The entire simulation methodology used here is GPU-accelerated and made freely available as part of the Castro simulation code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fast Machine Learning for Quantum Control of Microwave Qudits on Edge Hardware

Quantum optimal control is a promising approach to improve the accuracy of quantum gates, but it relies on complex algorithms to determine the best control settings. CPU or GPU-based approaches often have delays that are too long to be applied in practice. It is paramount to have systems with extremely low delays to quickly and with high fidelity adjust quantum hardware settings, where fidelity is defined as overlap with a target quantum state. Here, we utilize machine learning (ML) models to determine control-pulse parameters for preparing Selective Number-dependent Arbitrary Phase (SNAP) gates in microwave cavity qudits, which are multi-level quantum systems that serve as elementary computation units for quantum computing. The methodology involves data generation using classical optimization techniques, ML model development, design space exploration, and quantization for hardware implementation. Our results demonstrate the efficacy of the proposed approach, with optimized models achieving low gate trace infidelity near $10^{-3}$ and efficient utilization of programmable logic resources.

Sanders, Flor [Columbia U.]↗

Scaling the memory wall using mixed-precision - HPG-MxP on an exascale-class machine

Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.

Kashi, Aditya [ORNL] (ORCID:0000000325893792)↗