Search NASA⌕ Search

SEARCH · Search NASA

Results for “Massive parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Massive, long-lived electrostatic potentials in a rotating mirror plasma

Abstract Hot plasma is highly conductive in the direction parallel to a magnetic field. This often means that the electrical potential will be nearly constant along any given field line. When this is the case, the cross-field voltage drops in open-field-line magnetic confinement devices are limited by the tolerances of the solid materials wherever the field lines impinge on the plasma-facing components. To circumvent this voltage limitation, it is proposed to arrange large voltage drops in the interior of a device, but coexist with much smaller drops on the boundaries. To avoid prohibitively large dissipation requires both preventing substantial drift-flow shear within flux surfaces and preventing large parallel electric fields from driving large parallel currents. It is demonstrated here that both requirements can be met simultaneously, which opens up the possibility for magnetized plasma tolerating steady-state voltage drops far larger than what might be tolerated in material media.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

IMS Rapid Response 2024 Summary Report: A Machine Learning Potential for the Periodic Table

Stockpile stewardship and nuclear waste remediation are inherently chemically complex, involving practically the full diversity of the periodic table, but existing methods are too expensive or not functional for a large diversity of atom types. Overall, the field of machine learning interatomic potentials (MLIPs) has advanced dramatically in 2024 with large high-accuracy datasets existing for bulk, surface, and organic chemical systems and new online leaderboards for diverse chemistry. To participate in, and bring LANL interests into this ecosystem, here, we have built upon existing technologies created by LANL to create a framework capable of creating machine learning interatomic potentials (MLIPs) for over 90 atom types. Our results have created a massively diverse coordination complex training dataset more than 3 times the size of existing datasets, parallelized MLIP training over multiple GPUs, enabling the training of an MLIP spanning the periodic table at 20 times the speed of prior training on 32 GPUs. These advances are substantial towards creation on foundational MLIPs for LANL-specific application areas.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Hubble Space Telescope Observations within the Sphere of Influence of the Powerful Supermassive Black Hole in PKS 0745-191

We present Space Telescope Imaging Spectrograph observations from the Hubble Space Telescope of the supermassive black hole (SMBH) at the center of PKS 0745-191, a brightest cluster galaxy (BCG) undergoing powerful radio-mode active galactic nucleus (AGN) feedback (P cav ~ 5 × 10 45 erg s -1 ). These high-resolution data offer the first spatially resolved map of gas dynamics within an SMBH's sphere of influence under such powerful feedback. Our results reveal the presence of highly chaotic, nondrotational ionized gas flows on subkiloparsec scales, in contrast to the more coherent flows observed on larger scales. While radio-mode feedback effectively thermalizes hot gas in galaxy clusters on kiloparsec scales, within the core, the hot gas flow may decouple, leading to a reduction in angular momentum and supplying ionized gas through cooling, which could enhance accretion onto the SMBH. This process could, in turn, lead to a self-regulating feedback loop. Compared to other BCGs with weaker radio-mode feedback, where rotation is more stable, intense feedback may lead to more chaotic flows, indicating a stronger coupling between jet activity and gas dynamics. Additionally, we observe a sharp increase in velocity dispersion near the nucleus, consistent with a very massive M BH ~ 1.5 × 10 10 M ⊙ SMBH. The density profile of the ionized gas is also notably flat, paralleling the profiles observed in X-ray gas around galaxies where the Bondi radius is resolved. These results provide valuable insights into the complex mechanisms driving galaxy evolution, highlighting the intricate relationship between SMBH fueling and AGN feedback within the host galaxy.

79 ASTRONOMY AND ASTROPHYSICS↗

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗

Broadband unidirectional visible imaging using wafer-scale nano-fabrication of multi-layer diffractive optical processors

We present a broadband and polarization-insensitive unidirectional imager that operates at the visible part of the spectrum, where image formation occurs in one direction, while in the opposite direction, it is blocked. This approach is enabled by deep learning-driven diffractive optical design with wafer-scale nano-fabrication using high-purity fused silica to ensure optical transparency and thermal stability. Our design achieves unidirectional imaging across three visible wavelengths (covering red, green, and blue parts of the spectrum), and we experimentally validated this broadband unidirectional imager by creating high-fidelity images in the forward direction and generating weak, distorted output patterns in the backward direction, in alignment with our numerical simulations. This work demonstrates wafer-scale production of diffractive optical processors, featuring 16 levels of nanoscale phase features distributed across two axially aligned diffractive layers for visible unidirectional imaging. This approach facilitates mass-scale production of ~0.5 billion nanoscale phase features per wafer, supporting high-throughput manufacturing of hundreds to thousands of multi-layer diffractive processors suitable for large apertures and parallel processing of multiple tasks. Beyond broadband unidirectional imaging in the visible spectrum, this study establishes a pathway for artificial-intelligence-enabled diffractive optics with versatile applications, signaling a new era in optical device functionality with industrial-level, massively scalable fabrication.

36 MATERIALS SCIENCE↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

Overview of ASDEX upgrade results in view of ITER and DEMO

Experiments on ASDEX Upgrade (AUG) in 2021 and 2022 have addressed a number of critical issues for ITER and EU DEMO. A major objective of the AUG programme is to shed light on the underlying physics of confinement, stability, and plasma exhaust in order to allow reliable extrapolation of results obtained on present day machines to these reactor-grade devices. Concerning pedestal physics, the mitigation of edge localised modes (ELMs) using resonant magnetic perturbations (RMPs) was found to be consistent with a reduction of the linear peeling-ballooning stability threshold due to the helical deformation of the plasma. Conversely, ELM suppression by RMPs is ascribed to an increased pedestal transport that keeps the plasma away from this boundary. Candidates for this increased transport are locally enhanced turbulence and a locked magnetic island in the pedestal. The enhanced D-alpha (EDA) and quasi-continuous exhaust (QCE) regimes have been established as promising ELM-free scenarios. Here, the pressure gradient at the foot of the H-mode pedestal is reduced by a quasi-coherent mode, consistent with violation of the high-n ballooning mode stability limit there. This is suggestive that the EDA and QCE regimes have a common underlying physics origin. In the area of transport physics, full radius models for both L- and H-modes have been developed. These models predict energy confinement in AUG better than the commonly used global scaling laws, representing a large step towards the goal of predictive capability. A new momentum transport analysis framework has been developed that provides access to the intrinsic torque in the plasma core. In the field of exhaust, the X-Point Radiator (XPR), a cold and dense plasma region on closed flux surfaces close to the X-point, was described by an analytical model that provides an understanding of its formation as well as its stability, i.e., the conditions under which it transitions into a deleterious MARFE with the potential to result in a disruptive termination. With the XPR close to the divertor target, a new detached divertor concept, the compact radiative divertor, was developed. Here, the exhaust power is radiated before reaching the target, allowing close proximity of the X-point to the target. No limitations by the shallow field line angle due to the large flux expansion were observed, and sufficient compression of neutral density was demonstrated. With respect to the pumping of non-recycling impurities, the divertor enrichment was found to mainly depend on the ionisation energy of the impurity under consideration. In the area of MHD physics, analysis of the hot plasma core motion in sawtooth crashes showed good agreement with nonlinear 2-fluid simulations. This indicates that the fast reconnection observed in these events is adequately described including the pressure gradient and the electron inertia in the parallel Ohm’s law. Concerning disruption physics, a shattered pellet injection system was installed in collaboration with the ITER International Organisation. Thanks to the ability to vary the shard size distribution independently of the injection velocity, as well as its impurity admixture, it was possible to tailor the current quench rate, which is an important requirement for future large devices such as ITER. Progress was also made modelling the force reduction of VDEs induced by massive gas injection on AUG. The H-mode density limit was characterised in terms of safe operational space with a newly developed active feedback control method that allowed the stability boundary to be probed several times within a single discharge without inducing a disruptive termination. Regarding integrated operation scenarios, the role of density peaking in the confinement of the ITER baseline scenario (high plasma current) was clarified. The usual energy confinement scaling ITER98(p,y) does not capture this effect, but the more recent H20 scaling does, highlighting again the importance of developing adequate physics based models. Advanced tokamak scenarios, aiming at large non-inductive current fraction due to non-standard profiles of the safety factor in combination with high normalised plasma pressure were studied with a focus on their access conditions. A method to guide the approach of the targeted safety factor profiles was developed, and the conditions for achieving good confinement were clarified. Based on this, two types of advanced scenarios (‘hybrid’ and ‘elevated’

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja↗