Search NASA⌕ Search

SEARCH · Search NASA

Results for “GPU”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

ORCHA: A performance portability system for extreme heterogeneity

Heterogeneity is the prevalent trend in the rapidly evolving high-performance computing (HPC) landscape in both hardware and application software. The diversity in hardware platforms, currently comprising various accelerators and a future possibility of specializable chiplets, poses a significant challenge for scientific software developers aiming to harness optimal performance across different computing platforms while maintaining the quality of solutions when their applications are simultaneously growing more complex. Code synthesis and code generation can provide mechanisms to mitigate this challenge. We have developed a divide and conquer approach where different aspects of performance are handled by different stand-alone tools that are interfaced with the application through generated code. This portability system, ORCHA, enables users to configure and orchestrate their computations among available resources on a platform by specifying a high-level recipe, thereby permitting a many-to-many paradigm where each recipe results in a different variant of the application. The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system. Tools in ORCHA distribution are: CG-Kit for translating the recipe into an execution graph; Milhoja to execute the graph by orchestrating data and task movement among hardware resources; and Macroprocessor that enables users to define their own code-shorthand for higher composability and easier management of code variants. Additionally, the design of ORCHA permits tools to work in a plug-and-play mode where the application can build and run without CG-Kit and Milhoja, and either tool can be swapped out for other tools with similar capabilities by modifying the code generation portion of ORCHA. In this paper, we describe the design of ORCHA and the role that code-generation plays in isolating applications from tools. We demonstrate the breadth of configurations ORCHA enables with a case study in which an application configuration is realized on three distinct hardware mappings—a GPU-centric, a CPU/GPU balanced, and a CPU/GPU concurrent layouts by using different recipes.

Lee, Youngjun↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Road Lidar Dataset for the TxDOT Austin District

This is a road lidar data collection for developing road elevation models and road inundation mapping methodologies, a joint work between ORNL and The University of Texas at Austin. This dataset is generated as part of the flood transportation infrastructure, partly funded by the NOAA CIROH project. ORNL is a project partner for high-performance computing-empowered flood inundation mapping methodology R&D. The dataset is computed using a GPU-accelerated lidar data processing workflow developed at ORNL. The lidar data source is from TxGIO, the state lidar data collection site. The output dataset is in two formats: laz and copc. It is organized by TxDOT's maintenance sections, covering the Austin District. Data size: 3.86 billion road lidar points, 1.67% of the entire lidar data input Projection: EPSG:32614 (WGS84/UTM zone 14N) Website: https://web.corral.tacc.utexas.edu/nfiedata/road3d/austin_district/AustinMaintenanceSections_H_epsg6343_V_epsg5703/ LICENSE FOR USE -- MAPS AND DATA DISCLAIMER This resource is shared under the Creative Commons Attribution CC BY, http://creativecommons.org/licenses/by/4.0/ MAPS AND DATA DISCLAIMER The Oak Ridge National Laboratory (ORNL) shall not be held liable for improper or incorrect use of the data described or information contained on this map or associated series of maps. The data and related map graphics are not legal, land survey or engineering documents and are not intended to be used as such. ORNL gives no warranty, express or implied, as to the accuracy, reliability, utility or completeness of this information. The user of these maps and data assumes all responsibility and risk for the use of the maps and data. ORNL disclaims all warranties, representations or endorsements either express or implied, with regard to the information contained in this map product, including, but not limited to, all implied warranties of merchantability, fitness for a particular purpose or non-infringement. This preliminary map product is for research and review purposes only. It is not intended to be used for emergency management operational or life safety decisions at the local or regional governmental level or by the general public. Users requiring information regarding hazardous conditions or meteorological conditions for specific geographic areas should consult directly with their city or county emergency management office.

54 ENVIRONMENTAL SCIENCES↗

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES↗

VECTOR Phase 1 Dataset: CAV Trajectory and Energy Consumption Records

This dataset contains benchmark experimental data from Phase 1 of the VECTOR project, focusing on the energy impact of CAV hardware components. The dataset includes vehicle trajectory data (speed and position) and corresponding energy consumption records collected from a CAV platform equipped with lidar, cameras, onboard computation units, and communication modules. The primary objective is to quantify the baseline energy consumption attributable to sensing and computing systems, independent of any advanced cooperative control strategies. During experiments, the leading vehicle followed a predetermined velocity profile, and the following CAV mirrored this trajectory using a basic car-following control to ensure consistent driving behavior. This setup enables a reliable benchmark for assessing the energy cost introduced by onboard CDA hardware (e.g., lidar and GPU-based processing). The dataset is essential for evaluating energy baselines and supports future comparative studies involving additional cooperative strategies. ![system img](system.png) ![vector img](vector.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Machine learning at the Spallation Neutron Source accelerator and target

We describe the ongoing efforts to apply Machine Learning techniques to improve the performance of our accelerator and target. Specially, we are looking to minimize halo beam losses in the absence of a proper physics model, automatically detect and log anomalies in the target support systems such as cooling, and detect and prevent errant beam pulses in the linac. We also describe the infrastructure we use to acquire and stream data to the GPU cluster for training, our code development cycle, and edge computing for model inference. To minimize halo beam losses, we use a Reinforcement Learning technique tested on a virtual accelerator. The target anomaly detection is trained on archived data using incomplete physics models and is made part of the existing target reporting system. The errant beam prevention analyzes beam current and beam phase waveforms as well as accelerator configuration data to predict errant pulses. We also develop continual learning to adapt to changes in the accelerator.

Accelerator Physics↗

Singularity-EOS: Performance Portable Equations of State and Mixed Cell Closures

We present Singularity-EOS, a new performance-portable library for equations of state and related capabilities. Singularity-EOS provides a large set of analytic equations of state, such as the Gruneisen equation of state, and tabulated equation of state data under a unified interface. It also provides support capabilities around these equations of state, such as Python wrappers, solvers for finding pressure-temperature equilibrium between multiple equations of state, and a unique modifier framework, allowing the user to transform a base equation of state, for example by shifting or scaling the specific internal energy. All capabilities are performance portable, meaning they compile and run on both CPU and GPU for a wide variety of architectures.

97 MATHEMATICS AND COMPUTING↗

xesn: Echo state networks powered by Xarray and Dask

Xesn is a Python package that allows scientists to easily design Echo State Networks (ESNs) for forecasting problems. ESNs are a Recurrent Neural Network architecture introduced by Jaeger (2001) that are part of a class of techniques termed Reservoir Computing. One defining characteristic of these techniques is that all internal weights are determined by a handful of global, scalar parameters, thereby avoiding problems during backpropagation and reducing training time significantly. Because this architecture is conceptually simple, many scientists implement ESNs from scratch, leading to questions about computational performance. Xesn offers a straightforward, standard implementation of ESNs that operates efficiently on CPU and GPU hardware. The package leverages optimization tools to automate the parameter selection process, so that scientists can reduce the time finding a good architecture and focus on using ESNs for their domain application. Importantly, the package flexibly handles forecasting tasks for out-of-core, multi-dimensional datasets, eliminating the need to write parallel programming code. Xesn was initially developed to handle the problem of forecasting weather dynamics, and so it integrates naturally with Python packages that have become familiar to weather and climate scientists such as Xarray (Hoyer & Hamman, 2017). However, the software is ultimately general enough to be utilized in other domains where ESNs have been useful, such as in signal processing (Jaeger & Haas, 2004).

97 MATHEMATICS AND COMPUTING↗

High-Resolution Simulations of Geological CO 2 Injection: Application to the SPE11 Benchmark

Geological carbon sequestration (GCS) will play a critical role in decarbonization and in facilitating the transition to clean energy systems. Because CO 2 is highly mobile, ensuring its safe and permanent injection into subsurface geological formations involves monitoring over larger spatial domains and longer time periods than is typical for hydrocarbon reservoirs. This can benefit from simulation tools capable of modeling key CO 2 trapping mechanisms, particularly those optimized for speed and scalability on high-performance computing systems. Using isothermal versions of the SPE11B and SPE11C benchmark cases, we conduct a mesh refinement study simulating CO 2 injection into kilometer-scale rock formations at centimeter resolution with the GEOS open-source simulation framework. We focus on how mesh refinement improves the accuracy of convective mixing in both 2D and 3D simulations. The computational costs associated with achieving a converged solution highlight the need for predictive upscaling techniques. A systematic performance scaling analysis—including both central processing unit (CPU) and graphics processing unit (GPU) architectures—complements the “Results” section.

Geosciences↗

Center for Tokamak Transient Simulations (RPI Unstructured Mesh Developments for FES SciDAC4 Partnerships) (Final Report)

The goal of this project was the development of unstructured mesh technologies for fusion simulation codes” for FES SciDAC partnerships and to integrate those technologies into the simulation workflows of those partnerships. Specific developments were executed in support of the following FES SciDAC4 partnerships: Partnership Center for High‐fidelity Boundary Plasma Simulation (HBPS), Center for Integrated Simulation of Fusion Relevant RF (RF‐SciDAC), Center for Plasma Surface Interactions: Predicting the Performance and Impact of Dynamic PFC Surfaces (PSI2), and Center for Tokamak Transient Simulations (CTTS). The key unstructured mesh development areas addressed in this project include (i) methods to most effectively perform PIC calculations on unstructured meshes; (ii) creating meshes for fusion systems accounting for any desired level of geometric complexity and providing physics aware mesh configurations, (iii) adapting unstructured meshes to control the discretization errors, (iv) executing unstructured mesh calculations on GPU accelerated systems, (v) supporting physics/application‐specific PIC operations including surface/wall interactions of particles, (vi) providing infrastructure tools to support the interactions of solvers with unstructured meshes, and (vii) providing advanced methods for coupling plasma physics codes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Membrane Development for CO2 Capture from Steel Manufacturing

The U.S. government is targeting a net-zero carbon-emission economy by 2050, offering an exciting opportunity for membrane-based CO2 capture from various industrial point sources. Given that industrial flue gas has low CO2 partial pressures and high volumetric flow rates, high-permeance membranes are needed to make membrane technology economically viable for large-scale deployment. Thin film composite (TFC) membranes are necessary for this implementation because they can provide high permeance by forming a thin selective layer on top of a porous support. This presentation will report the rational design and fabrication of NETL’s highly permeable non-aging TFCs achieved by: synthesizing a high-performance rubbery selective material; developing a high-porosity membrane support; optimizing coating methods to assemble the two materials into scalable membranes; and scaling up membrane supports and TFCs via a roll-to-roll process. The novel rubbery selective material shows mixed-gas CO2 permeability of 930 Barrer and CO2/N2 selectivity of 44, exceeding the 2008 Robeson upper bound. The resulting TFCs yield remarkably high CO2 permeance of 4,500 GPU and CO2/N2 selectivity of 34 at 23ºC. Moreover, the TFCs exhibit not only excellent performance stability (or non-aging behavior) for 1,000 hours in the lab, but also maintain their separation properties in a 700-hour field test at the U.S. DOE’s National Carbon Capture Center using real humid flue gas.

Tran, Thien↗

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

CRADA Final Report: CRADA Number NFE-22-09330 with General Fusion

General Fusion is developing a magnetized target fusion (MTF) approach that involves compressing an initial magnetically confined plasma inside a cavity formed in liquid metal. This approach builds from concepts initially developed under the Linus program at the U.S. Naval Research Laboratory and combines it with advances from compact toroid experiment (CTX) and sustained spheromak plasma experiment (SSPX) in compact toroid plasmas and coaxial Marshall gun systems. Modeling the tokamak during compression is central to designing a successful MTF device. The plasma is formed by coaxial helicity injection in the General Fusion device. Immediately after formation, the plasma has a diverted tokamak configuration with a single null. As the wall moves inwards, the plasma is repelled from the conducting surface and driven inwards by currents induced by its magnetic field in the liquid metal wall. As the liquid metal closes (or bridges) the opening of the coaxial plasma injector, the magnetic field topology alters to remove the null. Due to this, the plasma moves from a diverted to a wall-limited configuration. The liquid metal liner continues to close in and change shape, reducing in radius by a factor of ten at the peak of plasma compression. A model of the MTF plasma must be able to handle this continually varying geometry, and to be predictive, it must faithfully include the real imperfections arising in the process. In this project, we pursued a Monte Carlo approach to closures for MHD by computing kinetic electron trajectories in an MHD plasma background from simulations of GF devices. This requires enhancing the capabilities of the KORC-T code for running large ensembles of kinetic trajectories by porting it to GPU architectures and enabling workflows for large ensembles on OLCF machines. With these capabilities, it is possible to produce a large library of kinetic calculations of electron orbits evolving in plasma configurations spanning the magnetic configurations and plasma density profiles, including non-axisymmetry, arising in the General Fusion’s existing PI3 spherical tokamak device. Using ensembles will capture particles passing a single point in space in a given magnetic configuration, and the entire dataset will cover a range of global magnetic field geometries. By sampling around many starting points, this dataset will capture the spatial dependence of the plasma parameters. From this large dataset, it is possible to produce a reduced model for the kinetic effects not captured in MHD.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Bringing randomized algorithms to mainstream numerical linear algebra

Numerical linear algebra (NLA) underpins huge swaths of computational science and engineering. For scientists and engineers to make the most of the DOE’s computing resources, it is essential that they have access to high-performance implementations of algorithms with best-in-class scalability and reliability. Despite this, prevailing NLA libraries have little to no support for breakthrough algorithms from the field of randomized numerical linear algebra (RandNLA) that have been developed over the past twenty years. The goal of this LDRD was to break a log-jam that had prevented broad adoption of RandNLA. Our work had two thrusts. The first was to develop RandBLAS: a trustworthy and high-performance C++ library for randomized dimension reduction (an operation widely known as sketching). The second was the development of a novel randomized algorithm for computing a challenging type of matrix decomposition known as Householder QR with column pivoting (Householder QRCP). In this one-year late-start LDRD we successfully delivered RandBLAS 1.0 and new CPU and GPU codes for Householder QRCP. RandBLAS has extensive documentation at https://randblas.readthedocs.io/en/stable/. Papers on RandBLAS and and our high-performance QRCP codes are forthcoming.

97 MATHEMATICS AND COMPUTING↗