Search NASA⌕ Search

SEARCH · Search NASA

Results for “Exascale”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems

Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronizat and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

exadigitUE5

This project provides the AR/VR interface to ORNL's exascale digital twin. The main functionality is implemented using Unreal Engine 5.1 for Desktop or Microsoft Hololens2 based visualization and interation with the system. The digital twin provides data ingestion from telemetry, as well as triggering and interacting with simulations developed for the wider ExaDigiT project at ORNL, as well as for the LUMI system at CSC and other CrayEX Supercomputers. For the overarching project, see ExaDigiT at https://exadigit.github.io, with the code repositories at https://code.ornl.gov/exadigit.

Maiterth, Matthias [Oak Ridge National Laboratory ↗

TChem-atm v1.0

SAND2024-11300O TChem-atm is a software library that was developed to solve complex kinetic models for atmospheric chemistry applications. TChem-atm interface employs a hierarchical parallelism design to exploit the massive parallelism available from modern computing platforms. It also supports gas atmospheric chemistry applications, e.g., the energy exascale earth system model. TChem can be used as a box model or coupled with a climate model to compute the time evolution of gas tracer species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

UFNet: Joint U-Net and Fully Connected Neural Network to Bias Correct Precipitation Predictions from

Paper information. Shuang Yu, Indrasis Chakraborty, Gemma J. Anderson, Donald D. Lucas, Yannic Lops, and Daniel Galea. UFNet: Joint U-Net and fully connected neural network to bias correct precipitation predictions from climate models. Artificial Intelligence for the Earth Systems, 2024. Overview. This work develops the UFNet methodology to correct E3SM historical precipitation projection bias. The UFNet deep learning framework consists of a two-part architecture: a U-Net convolutional network to capture the spatiotemporal distribution of precipitation and a fully connected network to capture the distribution of higher-order statistics. The joint network, termed UFNet, can simultaneously improve the spatial structure of the modeled precipitation and capture the distribution of extreme precipitation values. Below we provide guidance for applying UFNet to correct the Energy Exascale Earth System Model (E3SM; Golaz et al. 2019) daily precipitation projection over the contiguous United States (CONUS). Getting started 1. Obtain the historical climate simulation and observation data. The E3SM historical simulation data are available through https://aims2.llnl.gov/search/cmip6/. The CPC unified gauge-based analysis of daily precipitation can be found through https://psl.noaa.gov/data/gridded/data.cpc.globalprecip.html. The ECMWF atmospheric reanalysis of the 20th century (ERA-20C) data are available through https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-20c. The spatial resolution of E3SM and observed datasets are both regridded to a common 1° resolution grid using conservative interpolation. The regridded E3SM, CPC and ERA-20C with 1° resolution can be found throught ./data/. 2. Train the fully connected network (DNN) Python train_dnn.py 3. Train the UFNet Python train_ufnet.py 4. Evaluation and compared with the baseline Python evaluation.py

Lucas, DonaldD↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

NGEE Arctic Field-to-Model

This repository contains workshop documentation and shell scripts/computational infrastructure for building, running, and analyzing output from the DOE Energy Exascale Earth System Model (E3SM) Land Model (ELM).

Fiorella, Rich [@lanl]↗

ExaChem/exachem

Open Source Exascale Quantum Chemistry Software

Panyala, Ajay [Pacific Northwest National Laborato↗

Hacking Kilometer-Scale Models: A Participative Model for Climate Information

In May 2025, nearly 700 participants from all around the world coalesced at 10 regional nodes and a few satellite nodes to take part in a global hackathon of kilometer-scale (horizontal grid spacing < 10 km) regional and global Earth system models. Exciting science is emerging from these efforts, ranging across novel model analysis, new ways of integrating with satellite data, and emulation with machine learning. New technologies were trialed that enable the community to work in new and complementary ways to democratize access to global information at a local scale from a set of the world’s highest-resolution climate models. The hackathon demonstrated how exascale data can be organized to be accessible to anyone. Fundamentally, the community could apply these techniques and technologies to move toward more participative models for coproduction and delivery of diverse sources of climate information for climate scientists and citizens alike.

Climate models↗

Effects of Surface Turbulence Flux Parameterizations on the MJO: The Role of Ocean Surface Waves

This study investigates the sensitivity of the Madden–Julian oscillation (MJO) to changes to the bulk flux parameterization and the role of ocean surface waves in air–sea coupling using a fully coupled ocean–atmosphere–wave model. The atmospheric and ocean model components of the Energy Exascale Earth System Model (E3SM) are coupled to a spectral wave model, WAVEWATCH III (WW3). Two experiments with wind speed–dependent bulk algorithms (NCAR and COARE3.0a) and one experiment with wave-state-dependent flux (COR3.0a-WAV) were conducted. We modify COARE3.0a to include surface roughness calculated within WW3 and also account for the buffering effect of waves on the relative difference between air-side and ocean-side momentum flux. Differences in surface fluxes, primarily caused by discrepancies in drag coefficients, result in significant differences in MJO’s properties. While COARE3.0a has better convection–circulation coupling than NCAR, it exhibits anomalous MJO convection east of the date line. The wave-state-dependent flux (COR3.0-WAV) improves the MJO representation over the default COARE3.0 algorithm. Strong easterlies over the Pacific Ocean in COARE3.0a enhance the latent heat flux (LHFLX). This is responsible for the anomalous MJO propagation after the date line. In COR3.0a-WAV, waves reduce the anomalous easterlies, leading to a decrease in LHFLX and MJO dissipation after the date line. These findings highlight the role of surface fluxes in MJO simulation fidelity. Most importantly, we show that the proper treatment of wave-induced effects in bulk flux parameterization improves the simulation of coupled climate variability.

54 ENVIRONMENTAL SCIENCES↗

Extrapolar Cloud Feedbacks as a Driver of Arctic Amplification

The role of cloud feedbacks in Arctic amplification (AA) of anthropogenic warming remains unclear. Traditional feedback analysis diagnoses the net cloud feedback as strongly positive in the tropics but either weak or negative in the Arctic, suggesting that AA would be amplified if cloud feedbacks were suppressed. However, in cloud-locking experiments using the slab ocean version of the Energy Exascale Earth System Model (E3SM), we find that suppressing cloud feedbacks results in a substantial decrease in AA under greenhouse gas forcing. We show that the increase in AA from cloud feedbacks arises from two main mechanisms: 1) the additional energy contributed by positive cloud feedbacks in the tropics leads to increased poleward moist atmospheric heat transport (AHT) which then amplifies Arctic warming; and 2) the additional Arctic warming is amplified by positive noncloud feedbacks in the region, together making extrapolar cloud feedbacks amplify AA. We also find that cloud changes can modify the strength of noncloud feedback, but that modification has a small effect on Arctic warming. We further examine the role of cloud feedbacks in AA using a moist energy balance model, which demonstrates that interactions of cloud feedbacks with moist AHT and other positive feedbacks dominate the influence of clouds on the pattern of surface warming. However, the contribution of cloud-induced changes in noncloud feedbacks on AA is relatively minor. Furthermore, these results demonstrate that traditional attributions of AA, that are based on local feedback analysis, overlook key interactions between extrapolar cloud changes, poleward AHT, and noncloud feedbacks in the Arctic.

58 GEOSCIENCES↗

Examining Cloud Feedback Components in the Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM)

Cloud feedback remains the main source of uncertainty in climate sensitivity estimated by global climate models (GCMs), largely because subgrid cloud responses are parameterized in GCMs due to their coarse resolution. Here, this study examines cloud feedback in the global 3.25-km Simple Cloud-Resolving Energy Exascale Earth System Model (E3SM) Atmosphere Model (SCREAM 3 km) through a pair of 1-yr atmosphere-only simulations with control and +4-K sea surface temperature perturbations. SCREAM 3 km produces a positive cloud feedback that falls within but at the upper end of the range of Coupled Model Intercomparison Project phase 5 (CMIP5) and CMIP phase 6 (CMIP6) models and expert judgment. The positive cloud feedback arises from positive contributions from both high- and low-level clouds, with increases in high-cloud altitude and decreases in low-cloud amount and optical depth playing key roles. The stronger-than-CMIP-average feedback is mainly attributable to the high-cloud altitude feedback, owing to cloud tops rising nearly isothermally in SCREAM 3 km. The positive low-cloud amount feedback is weaker in SCREAM than in GCMs because estimated inversion strength (EIS) increases more dramatically with warming. A coarser 12-km resolution version of SCREAM exhibits a weaker positive cloud feedback than SCREAM 3 km, mainly because its low-cloud-radiative flux is more sensitive to EIS, leading to a stronger negative low-cloud amount feedback. With this process-level assessment of cloud feedback, this study reveals where SCREAM aligns with and diverges from conventional GCMs and expert assessment, providing insights to inform further model improvement and future expert assessment.

Cloud radiative effects↗

Quantifying the Impacts of Land-Cover Change on the Hydrologic Response to Hurricane Ida in the Lower Mississippi River Basin

Abstract The Lower Mississippi River basin (LMRB) has experienced significant changes in land cover and is one of the most vulnerable regions to hurricanes in the United States. Here, we study the impacts of land-cover change on the hydrologic response to Hurricane Ida in LMRB. By using an integrated surface–subsurface hydrologic model, Energy Exascale Earth System Model (E3SM) Land Model coupled with the three-dimensional ParFlow subsurface flow model (ELM-ParFlow), we simulate the effects of land-cover change on the flood volume and peak timing induced by rainfall from Hurricane Ida. The results show that land-cover changes from 1850 to 2015, which resulted in a smoother surface and less vegetation, exacerbated both flood peak time and volume induced by Hurricane Ida. The effects of land-cover changes can be decomposed into two mechanisms: a smoother surface routes more water faster to a watershed outlet and less vegetation allows more water to contribute to surface runoff. By comparing scenarios in which the two mechanisms were isolated, we found that changes in soil moisture due to vegetation cover change have more dominant effects on floods in the southern part and changes in Manning’s coefficient have the largest effect on floods in the northern part of the LMRB. The study provides important insights into the complex relationship between land-use, land-cover, and hydrologic processes in coastal regions.

54 ENVIRONMENTAL SCIENCES↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING↗

Parallelizing autotuning for HPC applications: Unveiling the potential of the speculation strategy in Bayesian optimization

In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.

Bayesian optimization↗

Understanding power and energy utilization in large scale production physics simulation codes

Power is an often-cited reason for the move to advanced architectures on the path to Exascale computing. Here, this is due to practical considerations related to delivering enough power to successfully site and operate these machines, as well as concerns about energy usage while running large simulations. Since obtaining accurate power measurements can be challenging, it may be tempting to use the processor thermal design power (TDP) as a surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advanced technology systems at Lawrence Livermore and Sandia National Labs, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale Lawrence Livermore simulation codes are significantly more efficient than a simple processor TDP model might suggest.

HPC↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗