Search NASA⌕ Search

SEARCH · Search NASA

Results for “Supercomputing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

LDRD FY25 Program Overview

As Lawrence Livermore National Laboratory’s (LLNL’s) Laboratory Directed Research and Development (LDRD) program enters its fifth decade of leading-edge research and development, its impact and importance have never been stronger. The program continues to advance strategic investments in pioneering science, technology, and engineering, ensuring LLNL will be ready to deliver on our mission as it evolves over the coming decades. Investing in LDRD research, and the people who perform this critical work, gives LLNL the ability to sustain our role as a leader in the Department of Energy and National Nuclear Security Administration enterprise. The LDRD program enables high-risk, high-payoff research that anticipates emerging threats and future mission needs. By nurturing the ingenuity of the Lab’s greatest asset, its people, LDRD funding advances not only our research but also grows and nurtures our workforce: engaging future innovators with student mentoring, challenging postdoctoral researchers to apply their skills to support national security, and strengthening the leadership skills of early career staff. This annual report documents how LDRD investments advance LLNL’s science, technology, and engineering across our mission space. To assess LDRD’s impact we track both short and long-term metrics such as peer-reviewed publications, number of students, or professional fellows. In addition to reviewing these metrics, I encourage you to delve deeper into the breadth of science and technology that illustrate the strategic value of this research portfolio. For instance, a recent exploratory research project used advanced manufacturing to construct miniaturized three-dimensional ion traps for a quantum computer with reduced quantum error rates to enable applications that address national security missions and support basic science. Another project has delved into studying detonation by examining deflagration to enhance the safety and security of the nuclear weapons stockpile. LDRD researchers are also deploying AI agents on two of the world’s most powerful supercomputers to automate and accelerate inertial confinement fusion experiments. Other teams are delivering more accurate optical constants to enable improved validation for aluminum to advance atomic and molecular physics models. LDRD-driven discoveries of how metals deform under extreme conditions strengthen our ability to model and design materials for demanding national security environments. National security challenges are increasingly complex and continuously evolving. LDRD focuses our most innovative science and technology on these challenges, ensuring the Laboratory is developing creative, forward-leaning solutions for our nation and the world. The following pages feature highlights of published scientific advances, patents, and honors that stem from LDRD investments. As you read this report, I hope you will understand how these investments position the Laboratory, and our partners, to meet the demands of the decades ahead.

36 MATERIALS SCIENCE↗

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

Benchmarking DAOS Filesystem on Aurora

We benchmark the DAOS filesystem on Argonne's Aurora supercomputer (127 nodes, 4,064 targets) using fio, IOR, mdtest, and IO500 to characterize I/O and metadata performance across the DFS API and DFuse+POSIX. Single-client fio shows POSIX bandwidth saturating at 1–2 MiB I/O sizes, with write-heavy workloads outperforming reads. Multi-node IOR shows DFS bandwidth scaling well up to ~32 tasks/node, with write latency growing faster than read latency. An 8-node IO500 evaluation shows DFS achieving ~5x higher bandwidth and ~190x higher IOPS than POSIX. Results indicate DAOS is well-suited to read-heavy workloads like AI training data loading, given appropriately sized transfers and concurrency.

George, Rebecca [College of William and Mary, Will↗

FullWave — A Full Wave Parallel Code for Modeling RF Fields in Hot Tokamak Plasma

FullWave is a computer code that simulates how radio-frequency (RF) waves travel and deposit energy in the hot plasma inside a fusion reactor. RF waves are used to heat the plasma and drive electrical current, which is essential for sustaining fusion reactions. The code uses a new algorithm that can handle much finer spatial detail than previous codes — more than 100 times finer — while running efficiently on national supercomputers. It incorporates a detailed physics model that captures subtle kinetic effects important for accurate prediction of wave behavior. Under this project, FullWave was extended to cover multiple RF frequency ranges relevant to present and future tokamaks, and validated against experimental parameters from the DIII-D tokamak at General Atomics. Results were published in peer-reviewed journal articles.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SIMBER: the Simula Berkeley Education and Research Collaboration (CRADA Final Report)

The SIMBER project is centered around advancing the state-of-the-art in the science of modelling of the human heart and leveraging the collaborations between leading research groups at the University of California Berkeley (UC Berkeley), the Lawrence Berkeley National Laboratory (Berkeley Lab), and Simula Research Laboratory (Simula). In particular, the following areas of expertise are shared and expanded through this project: “heart-on-chip” experimental systems from UC Berkeley, mathematical modelling of the heart from Simula, and supercomputing software from Simula and Berkeley Lab. This intersection of expertise is already enabling the development of novel tools and knowledge that produce more accurate models of the human heart in both health and disease, which in turn are leading to drug screening technologies that make cardiac drug development faster, cheaper and more humane.

Li, Xiaoye Sherry [Lawrence Berkeley National Labo↗

Unitary Qubit Lattice Algorithms for Plasma Physics

This final technical report summarizes research conducted under DOE Award DE-SC0021653 to develop unitary Quantum Lattice Algorithms for modeling electromagnetic wave propagation and scattering in complex media, including plasmas. The project developed and validated quantum-inspired formulations of Maxwell's equations that preserve unitary evolution and can be evaluated on classical high-performance computing systems while providing a foundation for future quantum-computing implementations. Major accomplishments include the development of two- and three-dimensional algorithms for electromagnetic scattering; scalable, distributed-memory implementations demonstrated on the Perlmutter supercomputer; formulations for nonlinear lossless fluid dynamics and cold, lossless, inhomogeneous magnetized plasmas; and an explicit quantum algorithm for a time-discretized Lorenz model. Simulations reproduced a range of characteristic wave phenomena, including transient effects that are not readily apparent in conventional frequency-domain studies, demonstrating the effectiveness of the proposed approach for modeling complex electromagnetic and plasma systems. The work establishes a unified theoretical and computational framework for quantum and quantum-inspired simulation and provides a foundation for future implementation on fault-tolerant quantum systems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Viskores: Integrating Parallel Scientific Visualization Research into Applications

Viskores is a scientific visualization library that is the primary deployment of such algorithms to the parallel accelerated processors of modern DOE supercomputers. In this paper, we review the capabilities provided by Viskores and how these capabilities are leveraged by other software in the high-performance computing ecosystem. We discuss the Viskores data representation and pay particular attention to array management. Through this array management we describe how data is adapted between Viskores and other software along with strategies for converting dynamic, polymorphic objects to static representations better suited to GPU processing. We conclude with several examples of Viskores integrating with high-performance software that is used in production today.

Moreland, Ken [ORNL] (ORCID:0000000270513288)↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Optimizing Metadata Exchange: Leveraging DAOS for ADIOS Metadata I/O

In HPC I/O middleware like the Adaptable I/O System (ADIOS) often mediates data transfers between applications. The metadata I/O generated by such systems often presents significant scaling and performance limitations. This work seeks improvement opportunities for metadata I/O by leveraging the DAOS storage systems, a recent storage system solution deployed on high-end systems such as the Aurora supercomputer. We investigate the tradeoffs and the design space for integrating I/O engines for the ADIOS middleware based on the different storage mechanisms supported by DAOS. We present a new DAOS-Array-ChunkSize-aligned engine which provides up to 2.3× improved performance than when using the existing DAOS-POSIX interface, without requiring any application modifications.

Venkatesh, Ranjan Sarpangala↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Protein Structure Inspired Discovery of a Novel Inducer of Anoikis in Human Melanoma

Drug discovery historically starts with an established function, either that of compounds or proteins. This can hamper discovery of novel therapeutics. As structure determines function, we hypothesized that unique 3D protein structures constitute primary data that can inform novel discovery. Using a computationally intensive physics-based analytical platform operating at supercomputing speeds, we probed a high-resolution protein X-ray crystallographic library developed by us. For each of the eight identified novel 3D structures, we analyzed binding of sixty million compounds. Top-ranking compounds were acquired and screened for efficacy against breast, prostate, colon, or lung cancer, and for toxicity on normal human bone marrow stem cells, both using eight-day colony formation assays. Effective and non-toxic compounds segregated to two pockets. One compound, Dxr2-017, exhibited selective anti-melanoma activity in the NCI-60 cell line screen. In eight-day assays, Dxr2-017 had an IC50 of 12 nM against melanoma cells, while concentrations over 2100-fold higher had minimal stem cell toxicity. Dxr2-017 induced anoikis, a unique form of programmed cell death in need of targeted therapeutics. Our findings demonstrate proof-of-concept that protein structures represent high-value primary data to support the discovery of novel acting therapeutics. This approach is widely applicable.

Oncology↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Cholla-MHD: An Exascale-capable Magnetohydrodynamic Extension to the Cholla Astrophysical Simulation Code

Abstract We present an extension of the massively parallel, GPU native, astrophysical hydrodynamics code Cholla to magnetohydrodynamics (MHD). Cholla solves the ideal MHD equations in their Eulerian form on a static Cartesian mesh utilizing the Van Leer + constrained transport integrator, the HLLD Riemann solver, and reconstruction methods at second and third order. Cholla’s MHD module can perform ≈260 million cell updates per GPU-second on an NVIDIA A100 while using the HLLD Riemann solver and second order reconstruction. The inherently parallel nature of GPUs combined with increased memory in new hardware allows Cholla’s MHD module to perform simulations with resolutions ∼500 3 cells on a single high-end GPU (e.g., an NVIDIA A100 with 80 GB of memory). We employ GPU direct Message Passing Interface to attain excellent weak scaling on the exascale supercomputer Frontier, while using 74,088 GPUs and simulating a total grid size of over 7.2 trillion cells. A suite of test problems highlights the accuracy of Cholla’s MHD module and demonstrates that zero magnetic divergence in solutions is maintained to round off error. We also present new testing and CI tools using GoogleTest, GitHub Actions, and Jenkins that have made development more robust and accurate and ensure reliability in the future.

Astronomy & Astrophysics↗

GenASiS: General Astrophysical Simulation System. II. Self-gravitating Baryonic Matter*

GenASiS (General Astrophysical Simulation System) is a code being developed initially and primarily, though not exclusively, for the simulation of core-collapse supernovae on the world's leading capability supercomputers. This paper---the second in a series---documents capabilities for Newtonian self-gravitating fluid dynamics, including tabulated microphysical equations of state treating nuclei and nuclear matter (`baryonic matter'). Computation of the gravitational potential of a spheroid, and simulation of the gravitational collapse of dust and of an ideal fluid, provide tests of self-gravitation against known solutions. In multidimensional computations of the adiabatic collapse, bounce, and explosion of spherically symmetric pre-supernova progenitors---which we propose become a standard benchmark for code comparisons---we find that the explosions are prompt and remain spherically symmetric (as expected), with an average shock expansion speed and total kinetic energy that are inversely correlated with the progenitor mass at the onset of collapse and the compactness parameter.

Cardall, Christian [ORNL] (ORCID:000000020086105X)↗

PCMS: Parallel Coupler For Multimodel Simulations

This paper presents the Parallel Coupler for Multimodel Simulations (PCMS), a new GPU accelerated generalized coupling framework for coupling simulation codes on leadership class supercomputers. PCMS includes distributed control and field mapping methods for up to five dimensions. For field mapping PCMS can utilize discretization and field information to accommodate physics constraints. PCMS is demonstrated with a coupling of the gyrokinetic microturbulence code XGC with a Monte Carlo neutral transport code DEGAS2 and with a 5D distribution function coupling of an energetic particle transport code (GNET) to a gyrokinetic microturbulence code (GTC). Weak scaling is also demonstrated on up to 2,080 GPUs of Frontier with a weak scaling efficiency of 85%.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Applying corrective machine learning in the E3SM atmosphere model in C++ (EAMxx)

The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM) is the newest addition to the family of earth system models capable of explicitly resolving convective systems. SCREAM is a kilometer-scale configuration of the advanced E3SM Atmosphere Model (EAMxx), designed for heterogeneous computing architectures. While the enhanced accuracy of kilometer-scale modeling offers significant benefits, it comes with a substantial computational cost, limiting feasible simulation durations to only a few years to a few decades, even on the fastest supercomputers. Machine learning presents an opportunity for scientists to achieve the high accuracy of storm-resolving models at a significantly reduced cost. Building on the previous success of applying corrective machine learning (ML) to the FV3GFS earth system model, this study explores the effects of implementing corrective-ML in EAMxx-SCREAM. We also address the computational challenges of integrating our implementation of corrective-ML, which is written in Python, with the C++/Kokkos EAMxx driver, as well as potential reasons why this approach has not proved as effective for EAMxx-SCREAM as for FV3GFS.

Environmental sciences↗