Search NASASearch

DOE OSTI · 3739679

Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

Abstract

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.

Keep this discovery

BibTeXRIS

Mcdaniel, Adam [ORNL] (ORCID:000000016926028X), Jantz, Michael R. [University of Tennessee, Knoxville (UTK)], Sharma, Ashesh [Hewlett Packard Enterprise], Martin, Steven [Hewlett Packard Enterprise], Abbott, Steve [Hewlett Packard Enterprise], Khandekar, Shreyas [HPE], Neth, Brandon [Hewlett Packard Enterprise], Alvarez, Bruno Villasenor [Advanced Micro Devices (AMD)], Kashi, Aditya [ORNL] (ORCID:0000000325893792), Elwasif, Wael [ORNL] (ORCID:0000000305541036), Hernandez Mendoza, Oscar [ORNL] (ORCID:0000000253806951). 2026-06-01. Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes. https://doi.org/10.23919/isc.2026.11520492

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Vidyut3d: A Gpu Accelerated Fluid Solver for Non-Equilibrium Plasmas on Adaptive Grids

We present the numerical methods, programming methodology, verification, and performance assessment of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures, in this work. Our plasma fluid model solves the coupled conservation equations for species transport, electrostatic Poisson and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive-grid/particle management library, AMReX, and is portable over widely available vendor specific GPU architectures. We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth-order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on capacitive discharges and atmospheric pressure streamer propagation. We demonstrate the use of our solver on two 3D simulation cases: an atmospheric streamer propagation in Ar-H2 mixtures and a low pressure twin electrode radio frequency reactor. Our performance studies on three different CPU+GPU architectures indicate approximately 150-400X speed-up using AMD and NVIDIA GPUs per time step compared to a single CPU core for a 4 million cell simulation with 15 species.

Sitaraman, Hariswaran

Dust Composition of Comet 81P/Wild 2 From JWST Spectroscopy Compared to Stardust’s Fine-Grained Materials and Gems-Rich IDPS

The Stardust Mission returned samples from the coma of Jupiter Family comet 81P/Wild 2 for detailed laboratory analyses. Here we present and discuss the best-fit thermal dust model for the coma dust of comet 81P/Wild 2 as observed by JWST. Comet 81P was observed through JWST GO 018xx, using NIRSpec (2.9–5.3µm, λ/∆λ≈1000) and MRS IFU (4.9–28.1µm, λ/∆λ≈3000) on 2023-03-20 UT and 2023-03-24 UT, respectively, at a heliocentric distance of 1.85 au and JWST distance of 1.43 au (phase angle of 32 degrees). The dust coma of 81P as revealed by JWST offers a salient compliment to laboratory studies of Stardust samples that typically are bigger than 2 µm and up to 60 µm in size.

D H Wooden

Battery Material Synthesis and Scalability using a 50L Taylor Vortex Reactor (Final CRADA Report)

Under this agreement, Laminar will loan Argonne a 50L Taylor Vortex Reactor (TVR) and provide mechanical troubleshooting guidance and consulting to ensure the successful setup of the pilot-scale synthesis process. The U.S. Department of Energy (DOE) will allocate funding for the labor and materials required for the study. To evaluate the physical and electrochemical properties of the materials produced by the 50L TVR, Argonne will perform comprehensive characterizations, including XRD, SEM, PSA, ICP, tap density, and coin half-cell testing. Throughout the collaboration, Argonne will provide feedback and recommendations for mechanical improvements to the reactor system. Furthermore, Argonne will credit Laminar as a collaborator in any presentations or publications resulting from data generated by the system. Laminar will retain no rights to experimental results or intellectual property generated through the experiments conducted with the 50L TVR at Argonne.

25 ENERGY STORAGE