Search NASASearch

SEARCH · Search NASA

Results for “parallel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale

Analysis of the impact of parallel magnetic fluctuations on linear gyrokinetic stability in NSTX-U and verification of gyro-fluid models

In this work, we use the CGYRO gyrokinetic code to analyze two L- and one H-mode discharges from the National Spherical Torus Experiment (NSTX) and NSTX-Upgrade (NSTX-U) selected due to their different mix of ion-scale driftwaves, ion temperature gradient (ITG) mode and trapped electron mode (TEM), and electromagnetic instabilities, kinetic ballooning mode (KBM), and micro-tearing mode (MTM) in the plasma core. It is found that the effect of parallel magnetic fluctuations is strongly destabilizing to the unstable KBMs compared to calculations with only perpendicular magnetic fluctuations. Two discharges have a mix of ITG/TEM and MTMs that are predicted to be dominant instability across the plasma radius. The parallel magnetic fluctuations are found to have little effect on the MTM stability but are destabilizing to ITG/TEM modes. To test the validity of the gyro-fluid linear stability codes TGLF and GFS at low aspect ratio, a database of linear growth rates has been created using the CGYRO gyrokinetic code. The database is comprised of various parameter scans around a standardized set of NSTX-U core parameters. It contains a group of electrostatic cases and an electromagnetic group that includes the effects of perpendicular and parallel magnetic fluctuations. Comparing the results from the GFS and TGLF models, we find that GFS exhibits the best agreement with the database of CGYRO linear growth rates. Comparing the model results for the electromagnetic scans shows that GFS captures the effects of parallel magnetic fluctuations accurately, while the TGLF model does not, as it lacks sufficient perpendicular energy resolution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Implementation of Pfirsch–Schlüter parallel flow in x-ray imaging crystal spectrometer inversion analysis

The x-ray imaging crystal spectrometer (XICS) tomographic inversion code for Wendelstein 7-X (W7-X) has been modified to consider the effects of parallel flows and has been applied to analyze measurements taken during recent experimental campaigns. Previous analysis neglected the effects of parallel flows due to the primarily perpendicular geometry of the sightlines and the small magnitude predicted by neoclassical theory. To reconsider these effects, the incompressibility condition for plasma flows is used to calculate the parallel Pfirsch–Schlüter flow component for the equilibrium configuration. By incorporating this condition along with the geometry of the sightlines—i.e. the fractional contributions of perpendicular and parallel flows—, an updated expression for the measured flow is used for the profile inversion. Application of this modified inversion code to data from W7-X shows that the magnitude of the radial electric field and the flux surface averaged perpendicular flow are reduced by approximately a factor of 2 and brought into better agreement with neoclassical predictions and charge exchange recombination spectroscopy measurements.

Pfirsch–Schlüter flows

Scalable Parallel Measurement of Individual Nitrogen-Vacancy Centers

The nitrogen-vacancy (NV) center in diamond is a solid-state spin defect that has been widely adopted for quantum sensing and quantum information processing applications. Typically, experiments are performed either with a single isolated NV center or with an unresolved ensemble of many NV centers, resulting in a trade-off between measurement speed and spatial resolution or control over individual defects. In this work, we introduce an experimental platform that bypasses this trade-off by addressing multiple optically resolved NV centers in parallel. We perform charge- and spin-state manipulations selectively on multiple NV centers from within a larger set, and we manipulate and measure the electronic spin states of over 100 NV centers in parallel. We show that the high signal-to-noise ratio of the measurements enables the detection of shot-to-shot pairwise correlations between the spin states of 108 NV centers, corresponding to the simultaneous measurement of 5778 unique correlation coefficients. We discuss how our platform can be scaled to parallel experiments with thousands of individually resolved NV centers. These results enable parallelized high-throughput sensing experiments that retain the spatial resolution of single defects and will, thereby, help to unlock advances in applications such as single-molecule NMR and characterization of integrated circuits. In addition, our approach to multiplexing provides a natural platform for the application of recently developed correlated sensing techniques.

NV centers

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Unorthodox parallelization for Bayesian quantum state estimation

Quantum state tomography (QST) allows for the reconstruction of quantum states through measurements and some inference technique under the assumption of repeated state preparations. Bayesian inference provides a promising platform to achieve both efficient QST and accurate uncertainty quantification, yet is generally plagued by the computational limitations associated with long Markov chains. In this work, we present a novel Bayesian QST approach that leverages modern distributed parallel computer architectures to efficiently sample a D-dimensional Hilbert space. Using a parallelized preconditioned Crank–Nicholson Metropolis–Hastings algorithm, we demonstrate our approach on simulated data and experimental results from IBM Quantum systems up to four qubits, showing significant speedups through parallelization. Although highly unorthodox in pooling independent Markov chains, our method proves remarkably practical, with validation ex post facto via diagnostics like the intrachain autocorrelation time. We conclude by discussing scalability to higher-dimensional systems, offering a path toward efficient and accurate Bayesian characterization of large quantum systems.

Bayesian inference

System-Level Efficiency Study of Modular DC/DC and Fuel Cell Stack Series–Parallel Configurations

This paper presents a system-level efficiency study of modular DC/DC converter and fuel cell stack configurations under series, parallel, and series–parallel connections. The investigation considers selected DC/DC converter topologies, including isolated and non-isolated architectures, boost, non-inverting buck-boost, and resonant converters, integrated with commercially available fuel cell stacks, including the Accelera FCE150, Ballard FCgen-HPS, and Toyota TFCM2. DC/DC converter modules are evaluated in modular configurations rated at 60 kW and 90 kW and two distinct output voltage ranges, specifically 580–730V and 780–930V, examining how different interconnection schemes impact overall system efficiency. The converters are evaluated under full load (100%), partial load (66%), and light load (33%) conditions, providing a comprehensive assessment of efficiency and operational characteristics across varying power demands. Approximately two-thousand efficiency data points are obtained from laboratory prototype–level component measurements and validated design evaluations across multiple converter topologies, modular configurations, voltage ranges, and load conditions, providing a robust dataset for comparative system-level efficiency analysis. The results highlight the effects of modularity and topology selection on system-level efficiency, offering a “playbook” framework for designers to select appropriate DC/DC converter arrangements and fuel cell stack connections for series, parallel, or hybrid configurations based on efficiency considerations.

DC/DC

Parallel-in-Time Solution of Scalar Nonlinear Conservation Laws

Here, we consider the parallel-in-time solution of scalar nonlinear conservation laws in one spatial dimension. The equations are discretized in space with a conservative finite-volume method using weighted essentially nonoscillatory (WENO) reconstructions, and in time with high-order explicit Runge–Kutta methods. The solution of the global, discretized space-time problem is sought via a nonlinear iteration that uses a novel linearization strategy in cases of nondifferentiable equations. Under certain choices of discretization and algorithmic parameters, the nonlinear iteration coincides with Newton’s method, although, more generally, it is a preconditioned residual correction scheme. At each nonlinear iteration, the linearized problem takes the form of a certain discretization of a linear conservation law over the space-time domain in question. An approximate parallel-in-time solution of the linearized problem is computed with a single multigrid reduction-in-time (MGRIT) iteration; however, any other effective parallel-in-time method could be used in its place. The MGRIT iteration employs a novel coarse-grid operator that is a modified conservative semi-Lagrangian discretization and generalizes those we have developed previously for nonconservative scalar linear hyperbolic problems. Numerical tests are performed for the inviscid Burgers and Buckley–Leverett equations. For many test problems, the solver converges in just a handful of iterations with a convergence rate independent of mesh resolution, including problems with (interacting) shocks and rarefactions.

97 MATHEMATICS AND COMPUTING

Parallel-in-Time Solution of Hyperbolic PDE Systems via Characteristic-Variable Block Preconditioning

We consider the parallel-in-time solution of both linear and nonlinear hyperbolic partial differential equation (PDE) systems in one spatial dimension. In the nonlinear setting, the discretized equations are solved with a preconditioned residual iteration based on a global linearization. The linear(ized) equation systems are approximately solved parallel-in-time using a block preconditioner applied in the characteristic variables of the underlying linear(ized) hyperbolic PDE. This change of variables is motivated by the observation that intervariable coupling between characteristic variables is weak, at least locally where spatio-temporal variations in the eigenvectors of the associated flux Jacobian are sufficiently small, while that between the original variables is not. For an ℓ-dimensional system of PDEs, applying the preconditioner consists of solving a sequence of ℓ scalar linear(ized)-advection-like problems, each associated with a different characteristic wave-speed in the underlying linear(ized) PDE. Furthermore, we approximately solve these linear advection problems using multigrid reduction-in-time (MGRIT); however, any other suitable parallel-in-time method could be used. Numerical examples are shown for the (linear) acoustics equations in heterogeneous media and for the (nonlinear) shallow water equations and Euler equations of gas dynamics with shocks and rarefactions. For many test problems, the solver converges in just a handful of iterations and with mesh-independent convergence rates.

97 MATHEMATICS AND COMPUTING

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol

Osmotic control of the spacing of parallel shear cracks in shale growing subcritically in geologic past

The geological genesis of natural cracks in sedimentary rocks such as shale is a problem that needs to be understood to improve the technology of hydraulic fracturing as well as deep sequestration of harmful fluids. Why are the vertical natural cracks roughly parallel and equidistant, and why is the spacing roughly 10 cm rather than 1 cm or 100 cm? Fracture mechanics of critical cracks cannot answer this question. Neither can the material heterogeneity. The growth of critical parallel cracks is impossible because the relative crack face displacements would immediately localize into one crack, leading to an earthquake. The cracks must have formed, on the tectonic time scale, by a slow growth of subcritical shear cracks governed by the Charles-Evans law. The idea advanced here is that what controls the crack spacing is the balance between the reduction, due to shear dilatancy, of the concentration of ions such as Na + and Cl - in each fracture process zone (PFZ), which decelerates the cracks, and the restoration of ion concentration by diffusion of ions from the space between the cracks into the FPZ. This diffusion of water is driven mainly by the osmotic pressure gradient, which offsets the deceleration and depends strongly on the crack spacing. A simple analytical solution of the steady state is rendered possible by approximating the ion concentration profiles between adjacent cracks by parabolic arcs. Applying this theory to Woodford shale yields the approximate crack spacing of 10 cm, which is realistic. Furthermore, the stability of unlimited parallel mode II frictional crack growth is proven by examining the second variation of the free energy. Water concentration drop in the FPZ due to shear dilatancy and its restoration by water diffusion from the inter-crack space have similar effect, although probably much weaker.

42 ENGINEERING

A parallel-kinetic-perpendicular-moment model for magnetised plasmas

We describe a new model for the study of weakly collisional, magnetised plasmas derived from exploiting the separation of the dynamics parallel and perpendicular to the magnetic field. This unique system of equations retains the particle dynamics parallel to the magnetic field while approximating the perpendicular dynamics through a spectral expansion in the perpendicular degrees of freedom, analogous to moment-based fluid approaches. In so doing, a hybrid approach is obtained that is computationally efficient enough to allow for larger-scale modelling of plasma systems while eliminating a source of difficulty in deriving fluid equations applicable to magnetised plasmas. We connect this system of equations to historical asymptotic models and discuss advantages and disadvantages of this approach, including the extension of this parallel-kinetic-perpendicular moment beyond the typical region of validity of these more traditional asymptotic models. This paper forms the first of a multi-part series on this new model, covering the theory and derivation, alongside demonstration benchmarks of this approach that include shocks and magnetic reconnection.

astrophysical plasmas

High-throughput synthesis of high-entropy alloys via parallelized electric field assisted sintering

Materials discovery and design is an expensive and time-consuming process, though necessary to advance many engineering fields. In this work, a novel tooling design is utilized in conjunction with electric field assisted sintering (EFAS) to effectively create a new high-throughput synthesis technique: parallelized EFAS. Through this technique, a wide range of material compositions and geometries can be synthesized in parallel as isolated samples or as part of contiguous arrays. Multiple tooling designs are explored to examine both the flexibility and limitations of the technique. A series of increasing complex alloys is produced simultaneously using in situ alloying, beginning with pure Ni and adding equimolar constituents up to the septenary high-entropy alloy AlCoCrCuFeMnNi. Microstructural characterization reveals each sample is effectively fully dense and chemically homogenous while exhibiting phases in agreement with CALPHAD predictions. Scalability of parallelized EFAS is then experimentally demonstrated and the implications for materials discovery and automation are discussed.

36 - MATERIALS SCIENCE

Understanding cold electron impact on parallel-propagating whistler chorus waves via moment-based quasilinear theory

Earth's magnetosphere hosts a wide range of collisionless particle populations that interact through various wave-particle processes. Among these, cold electrons, with energies below 100 eV, often dominate the plasma density but remain poorly characterized due to measurement challenges such as spacecraft charging and photoelectron contamination. Understanding the contribution of these cold populations to wave–particle interaction is of significant interest. Recent kinetic simulations identified a secondary drift-driven instability, in which parallel-propagating whistler-mode chorus waves excite oblique electrostatic whistler waves near the resonance cone and Bernstein-mode turbulence. These secondary modes enable a new channel of energy transfer from the parallel-propagating whistler wave to the cold electrons. In this work, we develop a moment-based quasilinear theory of the secondary instabilities to quantify such energy exchange. Our results show that these secondary instabilities persist for a wide range of parameters and, in many cases, lead to nearly complete damping of the primary wave. Such secondary instability might limit the amplitude of parallel-propagating whistler waves in Earth's magnetosphere and might explain why high-amplitude oblique whistler or electron Bernstein waves are rarely observed simultaneously with high-amplitude field-aligned whistler waves in the inner magnetosphere.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY