Search NASASearch

SEARCH · Search NASA

Results for “AMD”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Energy-Efficient Selective Removal of Metal Ions from Mining Influenced Waters (MIW) Using H-Bonded Organic-Inorganic Framework (HOIFS) (CRADA Final Report)

The work developed a hydrogen-bonded organic–inorganic framework (HOIF), specifically zinc imidazole salicylaldoxime supramolecule (ZIOS), for selective Cu removal and recovery from acidic mining-impacted waters (AMD), with emphasis on scalable synthesis, membrane integration, durability in real AMD (RAMD), mechanistic understanding, and recovery/regeneration pathways. This research breaks new ground in selective resource recovery from AMD waters – no other sorbents can operate reliably in this pH range. This project is unique in what it has delivered and is of general use to the public for projects relating to critical materials recovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Innovations Driven by Advanced Characterization to Strategize Critical Mineral Production and Beneficial Reuse from Fossil Energy Waste

Critical minerals (CM), such as rare earth elements (REE), cobalt, nickel, and lithium, have important uses in modern electronics and advanced manufacturing, yet are vulnerable to potential supply chain disruptions. Relatively abundant and readily available fossil energy (FE) wastes, such as coal combustion ash, acid mine drainage (AMD) and treatment solids (AMD solids), and Oil and Gas (O&G) drilling wastes (drill cuttings and produced waters) are under consideration as CM feedstocks. The National Energy Technology Laboratory (NETL) has studied CM resources for various FE wastes as part of the U.S. Department of Energy’s mission of bolstering the domestic CM supply, and makes the data available to the public on EDX at sites such as the NEWTS group. Advanced characterization utilizing synchrotron x-ray techniques coupled with laboratory extractions has been performed to identify CM hosting phases in these FE wastes to inform CM recoverability mechanisms. Novel methods to selectively recover CMs while co-producing other valuable byproducts have been developed. Successful examples discussed here include: (1) The identification of REE/Co/Ni/Sc binding and hosting phases in select FE waste (coal combustion ash and AMD solids), resulting in the development of a patented CM step-extraction process, (2) coupled production of functional sorbents from these extraction wastes and for CM recovery. A pilot-scale testing to evaluate the patent’s technical feasibility for extracting REE from coal ash on a barrel scale has been successfully performed. Additionally, (3) evaluation and measurements of brine geochemistry from U.S. O&G produced waters has informed a high Li recovery potential from Marcellus Shale produced water. NETL researchers have been developing tailored pre-treatment processes, an innovative and highly durable lithium sorbent, and geochemical model guided precipitation to accelerate Li production from the Marcellus Shale produced waters. These innovations driven by characterization are integral for maximizing and advancing the potential for CM recovery while offsetting the cost and environmental footprint for FE waste management.

critical mineral processing

Pilot-Scale Modular Research Facility for Acidic Water Pollution Cleanup and Domestic Production of Critical Minerals for National Security

A recent study by Penn State researchers revealed that Pennsylvania AMD streams originate from abandoned mines, with coal refuse piles of the lower Kittanning coal seam containing the most valuable heavy rare earth elements. Penn State has developed a three-stage process to recover these critical minerals, tested it for proof of concept, and secured a patent. Funded by the US DOE, a modular pilot-scale research and development unit has been designed and built to process 1,000 gallons per day of AMD from a site managed by the Pennsylvania Department of Environmental Protection (PA DEP). The system will selectively recover iron, aluminum, rare earth elements, and cobalt-nickel-manganese concentrates from AMD, followed by a proprietary downstream purification process. These operations aim to produce concentrates of critical minerals while treating acid mine drainage to meet environmental standards and evaluate different feedstocks. This presentation will describe the design, operation, and innovations of the proposed process and pilot facility, emphasizing its role in promoting sustainable recovery of critical minerals from legacy waste streams.

Pisupati, Sarma V [Center for Critical Minerals, T

Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications

We characterize the GPU energy usage of two widely adopted exascale-ready applications representing two classes of particle and mesh solvers: (i) QMCPACK, a quantum Monte Carlo package, and (ii) AMReX-Castro, an adaptive mesh astrophysical code. We analyze power, temperature, utilization, and energy traces from double-/single (mixed)-precision benchmarks on NVIDIA’s A100 and H100 and AMD’s MI250X GPUs using queries in NVML and rocm_smi_lib, respectively. We explore application-specific metrics to provide insights on energy vs. performance trade-offs. Our results suggest that mixed-precision energy savings range between 6–25% on QMCPACK and 45% on AMReX-Castro. Also, we found gaps in the AMD tooling used on Frontier GPUs that need to be understood, while query resolutions on NVML have little variability between 1 ms-1 s. Overall, application level knowledge is crucial to define energy-cost/science-benefit opportunities for the codesign of future supercomputer architectures in the post-Moore era.

Godoy, William [ORNL] (ORCID:0000000225905178)

Revealing the electronic structure of van der Waals antiferromagnetic NiPS 3 through synchrotron-based 𝜇-ARPES and alkali metal dosing

Antiferromagnetic NiPS 3 has recently emerged as a quantum material of considerable interest, thanks to the discovery of multiple new couplings involving electrons, spins, orbitals, phonons, and magnons. However, controversies and open questions persist concerning the fundamental origins of these couplings. A critical piece of information required to advance the understanding is the precise electronic band structure of NiPS 3 . Angle-resolved photoemission spectroscopy (ARPES), combined with alkali metal dosing (AMD), can enable us to directly observe the subtle electronic states that appear around the Fermi surface, offering valuable insights into the intriguing quantum properties and interplays of the examined material. Here, in this study, we present a comprehensive characterization and analysis of the band structure of van der Waals layered antiferromagnet NiPS 3 , leveraging state-of-the-art μ-ARPES measurements supported by density functional theory (DFT) calculations. Theoretical DFT results identify the orbital contributions to the observed bands, providing a precise understanding of the experimental ARPES data. Crucially, AMD enables the observation of conduction band and defect-related states above the valence band maximum in NiPS 3 . Furthermore, temperature dependent ARPES results across the Néel transition temperature of NiPS 3 reveal that the paramagnetic and antiferromagnetic phases have nearly identical band structures, underlining the highly localized character of Ni d states. These findings substantially deepen our understanding of the electronic properties of NiPS 3 and lay a vital foundation for exploring the intriguing quantum phenomena it exhibits.

Cao, Yifeng [Boston Univ., MA (United States); Law

Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.Multi

This work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora

Overview of NLR Automated Mobility District Implementation Research - Phases I, II, and III. The Convergence of Automation, Electrification, and On-Demand Services: Enabling Resilient Automated Mobility Districts

A research program by the National Renewable Energy Laboratory has been investigating the implementation prospects for fully automated passenger transport systems that are deployed to operate within dense urban settings, referred to as Automated Mobility Districts (AMDs). An AMD emphasizes the deployment of automated vehicles (AV) passenger transport services within a dense urban setting and other major activity centers with intense passenger origin-destination demand patterns, such as those found in large business districts, airports, and university and medical campuses. Phase I and II surveyed 10 early deployment sites and subsequently collected and evaluated the lessons learned from these early deployment sites, with particular attention to fleet operations, impacts of service reliability, and vehicle technology evolution as the field of companies was being progressively winnowed by the challenges of full automation.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Characterizing the Impact of GPU Power Management on an Exascale System

As GPU-accelerated high-performance computing (HPC) systems approach exascale performance, controlling energy consumption without compromising throughput is essential. Architectures such as the AMD MI250X-based Frontier supercomputer provide runtime mechanisms like frequency and power capping, enabling energy tuning without modifying application code. Although both target energy reduction, they operate via distinct hardware control paths and influence workloads differently. We present a comprehensive evaluation of these strategies on a leadership-class system using diverse HPC proxy applications representative of production workloads. Our study analyzes performance–energy trade-offs across multiple capping levels, node counts (1 and 32), and application profiles. Results show that frequency capping generally achieves higher energy efficiency and scalability, with gains of up to 13.2% without performance loss, while power capping is more effective for single-node runs or bursty GPU utilization. We also provide practical guidelines to help system administrators and users balance energy efficiency and performance in large-scale scientific workloads.

Costa, Mariana [Universidade Federal do Rio Grande

Vidyut3d: A GPU accelerated fluid solver for non-equilibrium plasmas on adaptive grids

We present the numerical methods, programming methodology, verification, and performance assessment of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures, in this work. Our plasma fluid model solves the coupled conservation equations for species transport, electrostatic Poisson and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive-grid/particle management library, AMReX, and is portable over widely available vendor specific GPU architectures. We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth-order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on capacitive discharges and atmospheric pressure streamer propagation. We demonstrate the use of our solver on two 3D simulation cases: an atmospheric streamer propagation in Ar-H2 mixtures and a low pressure three-electrode radio frequency reactor. Our performance studies on three different CPU+GPU architectures indicate ~ 150-400X speed-up using AMD and NVIDIA GPUs per time step compared to a single CPU core for a 4 million cell simulation with 15 species.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Ammonium-coordinated exchanger (ACE) functionalized silica sorbents for recovering/removing aqueous anionic contaminants

There are limited studies of functionalized silica anion exchange sorbents used for critical/heavy metal recovery/removal relative to polymeric and other inorganic materials. This work features ammonium-coordinated exchanger (ACE) anion exchange particle sorbents prepared by either acid-washing epoxy-crosslinked polyethylenimine (PEI) hydrogen bonded within/to a silica particle sorbent (two-step method) or reacting a di-chlorinated crosslinker, α,α-dichloro-p-xylene (DPX), with PEI within silica (single-step method). Energy dispersive X-ray spectroscopy (EDS) and infrared spectroscopy confirmed the presence of -NH 2 + ···Cl - and -NH 3 + ···Cl - groups, which removed oxyanionic species –arsenate, selenate, chromate, sulfate, phosphate, and nitrate– plus bromide from ideal solutions, authentic acid mine drainage (AMD), and authentic flue gas desulfurization (FGD) wastewater. Affinity of the anions for ACE varied across single- and mixed-element solutions. However, affinity for CrO 4 2- was among the highest in both cases. Total anion uptake by the optimized ACE, PEI-E3-HCl_1.1, reached 1.2 mmol anion/g-sorb. (0.56 mmol CrO 4 2- /g), or ∼2.3 mmol negative charge/g-sorb. This was close to the 0.52 mmol CrO 4 2- /g of a commercial anion exchange resin. Near-consistent removal of 20–80 % of each anion from FGD during an eight-cycle adsorption-desorption (1 M NaCl) test predicted good ACE viability for testing under practical conditions at larger scale.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Delving into the depths of NGC 3783 with XRISM V. Broad-band modelling of ionized outflows

The Seyfert 1 galaxy NGC 3783 hosts a multi-phase warm absorber (WA) that has been extensively studied in the X-ray band. High-resolution spectra from 2000−2001 revealed a complex outflow with multiple ionization and velocity components. Two decades later, new XMM-Newton and XRISM observations allow us to investigate the long-term evolution of these outflows. We performed joint spectral modelling of the XMM-Newton /RGS and XRISM/Resolve time-averaged spectra using the pion photoionization code within SPEX. We derived the ionization parameter, column density, turbulent velocity, and outflow velocity for each absorption component, and investigated their thermal stability and absorption measure distribution (AMD) to characterize the physical and dynamical properties of the WA in NGC 3783 in 2024. We compared these results with the 2000−2001 epoch to assess long-term variability, stability, and possible changes in the absorber population. We identify eight WA components spanning log ξ = 1.08 − 3.38 and outflow velocities of 480−1230 km s −1 . The ranges of column densities and turbulent velocities remain broadly consistent with the WAs from 2000−2001, but the earlier data contained more low-ionization, high-velocity components. The total column density of all the outflows in 2024 is 1.5 times larger than in 2000−2001, which means that it has been replenished by fresh material. The dominant unresolved transition array (UTA) absorber (component B3) has increased its column density by a factor of three while maintaining a similar ionization parameter. The WAs in NGC 3783 have undergone significant structural and dynamical evolution over the past 24 years.

79 ASTRONOMY AND ASTROPHYSICS

Experience with the alpaka performance portability library in the CMS software

ion Library for Parallel Kernel Acceleration) is a header-only C++ library that provides performance portability across different back-ends, abstracting the underlying levels of parallelism. It supports serial and parallel execution on CPUs, and extremely parallel execution on NVIDIA, AMD and Intel GPUs.This contribution will show how alpaka is used in the CMS software to develop and maintain a single code base; to use different toolchains to build the code for each supported back-end, and link them into a single application; to seamlessly select the best backend at runtime, and implement portable reconstruction algorithms that run efficiently on CPUs and GPUs from different vendors. It will describe the validation and deployment of the alpaka-based implementation in the CMS High Level Trigger, and highlight how it achieves near-native performance.

Alawieh, Jaafar [CERN]

Automated Hybrid Variance Reduction on Advanced Architectures in the Shift Monte Carlo Code

Monte Carlo transport methods are the most accurate schemes for solving problems with complex energy and spatial features, but they come with a high computational cost. Although hybrid methods have enabled the use of Monte Carlo transport for a large class of problems, they still require significant computing resources. Modern multicore CPUs with large numbers of compute cores and graphical processing units (GPUs) provide opportunities to optimize the memory and run-time costs of hybrid Monte Carlo methods. This paper documents the development and analysis of three Monte Carlo transport algorithms that support hybrid transport using the consistent adjoint-driven importance sampling (CADIS) and forward-weighted CADIS methods in the Shift Monte Carlo code: history-based transport using static and dynamic threading on multicore CPUs and event-based transport enabling weight window tracking on GPUs. The results are shown for two challenging hybrid problems on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. The results show that all three methods yield good performance and enable solutions of difficult fixed-source transport problems in less than 2 min on 20 nodes of Frontier. Dynamic threading was observed to give up to 20% better scaling behavior than static threading. Moreover, the AMD Instinct 250X GPU was found to give 9 to 11 times greater throughput per graphics compute die than the best CPU performance. In conclusion, additional opportunities for optimization of hybrid transport on GPUs are discussed.

Denovo

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S

JACC.shared: Leveraging HPC Metaprogramming and Performance Portability for Computations That Use Shared Memory GPUs

In this work, we present JACC.shared, a new feature of Julia for ACCelerators (JACC), which is the performanceportable and metaprogramming model of the just-in-time and LLVM-based Julia language. This new feature allows JACC applications to leverage the high-performance computing (HPC) capabilities of high-bandwidth, on-chip GPU memory. Historically, exploiting high-bandwidth, shared-memory GPUs has not been a priority for high-level programming solutions. JACC.shared covers that gap for the first time, thereby providing a highlevel, portable, and easy-to-use solution for programmers to exploit this memory and supporting all current major accelerator architectures. Well-known HPC and AI workloads, such as multi/hyperspectral imaging and AI convolutions, have been used to evaluate JACC.shared on two exascale GPU architectures hosted by some of the most powerful US Department of Energy supercomputers: Perlmutter (NVIDIA A100) and Frontier (AMD MI250X). The performance evaluation reports speedup of up to 3.5× by adding only one line of code to the base codes, thus providing important accelerators in a simple, portable, and transparent way and elevating the programming productivity and performance-portability capabilities for Julia/JACC HPC, AI, and scientific applications.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (