Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

DC Microgrid Reliability Enhancement with Adaptive Converter Thermal Management

Due to the different device selections, aging levels, and thermal dissipation performance, some converters may take additional thermal stress on switching devices than others in paralleled converter systems, which will reduce system reliability. To address this problem, this paper proposes a power-sharing strategy with adaptive thermal management. First, the temperature-based power loss model and electrical-thermal model are established. Based on that, a high-accuracy IGBT junction temperature estimate considering the power loss-temperature coupling can be achieved. Further, the thermal-sharing for all the switching devices in paralleled converters can be achieved with the proposed adaptive thermal management strategy. The proposed strategy can change the power-sharing ratio adaptively according to the system operation conditions, which will contribute to the system reliability enhancement. The effectiveness of the proposed strategy is verified through PLECS thermal simulation and joint real-time simulation with Dspace and RT-box.

DC microgrid↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

Virtual Self-Excited Induction Generator-Based Grid-Forming Inverter Control for Robust Voltage Regulation Under Nonideal Loading

This paper presents a generator-inspired control methodology for grid-forming (GFM) inverters that deliberately emulates a self-excited induction generator so that the inverter can hold its voltage and frequency under difficult loading and severe terminal disturbances across wide voltage and frequency ranges. The design integrates a Lyapunov energy function-based inner loop to provide high bandwidth and strong disturbance rejection, and it complements this with a passivity-based argument that furnishes a coherent large-signal stability guarantee beyond small-signal limits. Analytical insights are developed via the Krylov-Bogoliubov-Mitropolsky averaging method, which reveals an intrinsic resistive droop characteristic; these closed-form relations both explain the observed dynamics and yield simple, decentralized tuning rules. The methodology is validated on a controller-hardware-in-the-loop platform and exercised in real time across balanced, unbalanced, and nonlinear loads, as well as during parallel operation. Across these scenarios, the inverter maintains balanced three-phase voltages, limits harmonic content, settles quickly with well-damped transients, and remains resilient when multiple units operate in parallel. The contributions are a self-excited-machine-inspired GFM controller with enhanced dynamic performance and robustness, a single stability rationale grounded in passivity, closed-form expressions that guide tuning, and comprehensive hardware-in-the-loop validations demonstrating effectiveness and superiority under challenging operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Probabilistic Visualization of Local Divergence of 2D Vector Fields with Independent Gaussian Uncertainty

This work focuses on visualizing uncertainty of local divergence of two-dimensional vector fields. Divergence is one of the fundamental attributes of fluid flows, as it can help domain scientists analyze potential positions of sources (positive divergence) and sinks (negative divergence) in the flow. However, uncertainty inherent in vector field data can lead to erroneous divergence computations, adversely impacting downstream analysis. While Monte Carlo (MC) sampling is a classical approach for estimating divergence uncertainty, it suffers from slow convergence and poor scalability with increasing data size and sample counts. Thus, we present a two-fold contribution that tackles the challenges of slow convergence and limited scalability of the MC approach. (1) We derive a closed-form approach for highly efficient and accurate uncertainty visualization of local divergence, assuming independently Gaussian-distributed vector uncertainties. (2) We further integrate our approach into Viskores, a platform-portable parallel library, to accelerate uncertainty visualization. In our results, we demonstrate significantly enhanced efficiency and accuracy of our serial analytical (speed-up up to 1946×) and parallel Viskores (speed-up up to 19698×) algorithms over the classical serial MC approach. We also demonstrate qualitative improvements of our probabilistic divergence visualizations over traditional mean-field visualization, which disregards uncertainty. We validate the accuracy and efficiency of our methods on wind forecast and ocean simulation datasets.

Ouermi, Timbwaoga [University of Utah]↗

Bias-Modulated ALD of ZnO: Insights into Precursor-Surface Interactions for ZnO Films

Atomic layer deposition (ALD) is widely used to deposit conformal thin films but is often limited in the tunability of the resulting material’s properties. Substrate bias and electric fields alter precursor-surface interactions and provide means to tune material properties. To explore this, we performed zinc oxide (ZnO) ALD using diethylzinc (DEZ) and water on silicon native oxide substrates at 150 °C in a sample holder designed to create a static electrical field by biasing one plate of a parallel plate capacitor-style sample holder during deposition. ZnO films prepared in an electric field/on a biased sample holder were thinner, changed relative crystalline composition, and contained more carbon compared to samples grown in identical sample holders without bias. The thickness was independent of the magnitude of the eletric field between plates, indicating that the primary driver for the change was substrate biasing not the electric field between plates of the parallel plate capacitor-style sample holder. Density functional theory calculations showed enhanced electron migration between dissociatively adsorbed DEZ molecules and the ZnO (002) facet with increasing force from an electric field at the substrate surface, which strengthens the electronic interactions between the surface and the adsorbate. These models offer a compelling explanation for the inhibited growth, changes in the crystallinity, and increase in carbon content of films grown in an electric field/on biased plates.

Jones, Jessica C. (ORCID:0000000174754620)↗

Femtojoule optical nonlinearity for deep learning with incoherent illumination

Optical neural networks (ONNs) are a promising computational alternative for deep learning due to their inherent massive parallelism for linear operations. However, the development of energy-efficient and highly parallel optical nonlinearities, a critical component in ONNs, remains an outstanding challenge. Here, we introduce a nonlinear optical microdevice array (NOMA) compatible with incoherent illumination by integrating the liquid crystal cell with silicon photodiodes at the single-pixel level. We fabricate NOMA with more than half a million pixels, each functioning as an optical analog of the rectified linear unit at ultralow switching energy down to 100 femtojoules per pixel. With NOMA, we demonstrate an optical multilayer neural network. Our work holds promise for large-scale and low-power deep ONNs, computer vision, and real-time optical image processing.

36 MATERIALS SCIENCE↗

Hybridization capture sequencing for Vibrio spp. and associated virulence factors

ABSTRACT Proliferation ofVibriospp. in aquatic ecosystems is associated with climate change and, concomitantly, increased incidence of vibriosis. They are autochthonous to aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing (HCS) was employed to profile low-abundanceVibriospp. in environmental samples. The HCS panel targeted a family of molecular chaperones (CPN60) specific to 69Vibriospp. and 162Vibrio-specific virulence factors. This approach was evaluated in parallel with traditional whole-community shotgun sequencing in a metagenomic analysis of water and oyster samples collected from the Chesapeake Bay. In addition,Vibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples were subjected to whole-genome sequencing to determine the genetic characteristics of pathogenicVibriospp. circulating in an aquatic environment. HCS, employed to determine the incidence and characterization of specificVibriospp., yielded significantly greater metagenomic insight, notably a variety of otherVibriospp., including detection ofVibrio cholerae,Vibrio fluvialis, andVibrio aestuarianus, in addition toVibrio parahaemolyticusandVibrio vulnificus, and also important virulence factors not detectable using traditional molecular methods. Thus, pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood. It is concluded that environmental surveillance should include HCS, a valuable tool for the detection and characterization of pathogenic agents in aquatic ecosystems, notably vibrios. IMPORTANCE The increasing prevalence of pathogenicVibriospp. in aquatic ecosystems, driven by climate change, is closely linked to a rise in cholera and vibriosis cases, emphasizing the need for improved environmental surveillance. Vibrios are naturally occurring in aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing was employed to profile low-abundanceVibriospp. in metagenomic samples, namely water and oysters collected from the Chesapeake Bay. This approach was evaluated in parallel with traditional whole-community shotgun sequencing and whole-genome sequencing ofVibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples. Results suggest pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood, when multiple methods are considered for environmental surveillance.

Microbiology↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources—such as varying physical groundings or data acquisition systems—and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Real-Time Bayesian Inference at Extreme Scale: A Digital Twin for Tsunami Early Warning Applied to the Cascadia Subduction Zone

We present a Bayesian inversion-based digital twin that employs acoustic pressure data from seafloor sensors, along with 3D coupled acoustic–gravity wave equations, to infer earthquake-induced spatiotemporal seafloor motion in real time and forecast tsunami propagation toward coastlines for early warning with quantified uncertainties. Our target is the Cascadia subduction zone, with one billion parameters. Computing the posterior mean alone would require 50 years on a 512 GPU machine. Instead, exploiting the shift invariance of the parameter-to-observable map and devising novel parallel algorithms, we induce a fast offline–online decomposition. The offline component requires just one adjoint wave propagation per sensor; using MFEM, we scale this part of the computation to the full El Capitan system (43,520 GPUs) with 92% weak parallel efficiency. Moreover, given real-time data, the online component exactly solves the Bayesian inverse and forecasting problems in 0.2 seconds on a modest GPU system, a ten-billion-fold speedup.

97 MATHEMATICS AND COMPUTING↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

TChem-atm v1.0

SAND2024-11300O TChem-atm is a software library that was developed to solve complex kinetic models for atmospheric chemistry applications. TChem-atm interface employs a hierarchical parallelism design to exploit the massive parallelism available from modern computing platforms. It also supports gas atmospheric chemistry applications, e.g., the energy exascale earth system model. TChem can be used as a box model or coupled with a climate model to compute the time evolution of gas tracer species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

Vidyut3d: A Non-Equilibrium Plasma Modeling Tool [SWR-24-101]

Vidyut3d is a massively-parallel plasma-fluid solver for low-temperature plasmas (LTPs) that supports both local field (LFA) and local mean energy (LMEA) approximations, as well as complex gas and surface-phase chemistry. The solver supports 2D and 3D domains, and uses AMReX's adaptive mesh refinement capabilities to increase the grid resolution around complex structures (e.g. streamer heads and sheaths) while maintaining a tractable problem size. Vidyut specializes in simulating various types of gas-phase discharges, as well as plasma/surface interactions and surface chemistry (e.g. for plasma-mediated catalysis applications). The solver also supports hybrid CPU/GPU parallelization strategies, and has demonstrated excellent scaling on various HPC architectures for problem sizes consisting of O(100 M) control volumes.

Sitaraman, Hariswaran↗

Extended flexibility of neutron Larmor diffraction for increased diffraction range in mosaicity measurements

Larmor diffraction (LD) is a neutron scattering technique that offers enhanced resolution by harnessing the Larmor precession of neutron spins in a magnetic field. By encoding subtle changes in neutron momentum transfer into significant alterations in the Larmor phase of neutron spins, LD can be employed to measure lattice expansion, lattice distortion, and mosaicity with exceptional resolution. As originally proposed by Rekveldt et al., LD necessitates the magnetic field boundaries to be tilted to be parallel to the crystal plane of interest. Here, this report explores the fundamental principles shared between LD and spin-echo small-angle neutron scattering (SESANS). Drawing inspiration from the flexibility of SESANS to adjust magnetic field boundaries to optimize the resolution in the measurement of the neutron momentum transfers q, we will demonstrate that the strict requirement of parallel alignment between the magnetic field boundaries and the crystal plane can be relaxed for the measurements of mosaicity. Such relaxation will expand the accessible diffraction angles for these situations that are highly constrained.

Larmor diffraction↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗