Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

A Partitioned - Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a platform for multiscale modeling of composites which can perform analysis of materials with any arbitrary number of length scales. The platform supports modularity, scalability, and interoperability using recursive procedures and data structures. A Macro solver driven parallelization scheme often limits the capability of NASMAT to scale as it has access to limited memory and number of cores (often one core/thread) and often forces to implement macro solver specific changes to the platform. In this work, a partitioned task-parallel approach is adopted, where the parallelization strategy adopted for NASMAT is independent of the macro solver and the computational resources are managed independently. The programming architecture takes into account the hierarchy of multiple scales (task-dependence) and the heterogeneous nature (dynamic load balancing) of computation through implementation of a hierarchy-informed task parallel model. The partitioned nature of the framework further extends the “plug and play” capability of NASMAT. preCICE, an open-source library for coupling multiphysics solver in a partitioned manner, is adopted to integrate NASMAT with an external macro solver by implementing a NASMAT adapter for preCICE. Speedup and scalability of the framework is studied for micromechanical models of varying size.

task-parallel↗

A Partitioned -Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a “plug and play” software package that allows users to conduct massively multiscale modeling of hierarchical and nonlinear materials. This work extends the scalability and improves the High Performance Computing friendliness of NASMAT by adopting a Partitioned Task-Parallel approach. Interoperability of NASMAT with external software is enhanced through preCICE, a open source library for multiphysics coupling in a partitioned manner. Enhancement through preCICE allows for easy integration of NASMAT to other macro solvers and dissociates the parallelization strategy adopted within NASMAT from the macro solver. The task-parallel framework based on Master-Worker approach is implemented as the parallelization scheme. The scheme accounts for hierarchy of multiple scales (task-dependence) and heterogeneous nature (dynamic load balancing) of computations. The applicability and scalability of the framework will be evaluated by analyzing large scale engineering problems through massively multiscale methods.

NASMAT↗

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

Kinetic study of shock formation and particle acceleration in laser-driven quasi-parallel magnetized collisionless shocks

Quasi-parallel magnetized collisionless shocks are believed to be one of the most efficient accelerators in the universe. Compared to quasi-perpendicular shocks, quasi-parallel shocks are more difficult to form in the laboratory and to simulate because of their large spatial scales and long formation times. Our two-dimensional particle-in-cell simulations show that the early stages of quasi-parallel shock formation are achievable in experiments planned for the National Ignition Facility and that particles accelerated by diffusive shock acceleration (DSA) are expected to be observable in the experiment. Repetitive ion acceleration by crossings of the shock front, a key feature of DSA, is seen in the simulations. Other characteristic features of quasi-parallel shocks such as upstream wave excitation by energetic ions are also observed, and energy partition between the ions and the electrons in the downstream of the shock is briefly discussed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analysis of the impact of parallel magnetic fluctuations on linear gyrokinetic stability in NSTX-U and verification of gyro-fluid models

In this work, we use the CGYRO gyrokinetic code to analyze two L- and one H-mode discharges from the National Spherical Torus Experiment (NSTX) and NSTX-Upgrade (NSTX-U) selected due to their different mix of ion-scale driftwaves, ion temperature gradient (ITG) mode and trapped electron mode (TEM), and electromagnetic instabilities, kinetic ballooning mode (KBM), and micro-tearing mode (MTM) in the plasma core. It is found that the effect of parallel magnetic fluctuations is strongly destabilizing to the unstable KBMs compared to calculations with only perpendicular magnetic fluctuations. Two discharges have a mix of ITG/TEM and MTMs that are predicted to be dominant instability across the plasma radius. The parallel magnetic fluctuations are found to have little effect on the MTM stability but are destabilizing to ITG/TEM modes. To test the validity of the gyro-fluid linear stability codes TGLF and GFS at low aspect ratio, a database of linear growth rates has been created using the CGYRO gyrokinetic code. The database is comprised of various parameter scans around a standardized set of NSTX-U core parameters. It contains a group of electrostatic cases and an electromagnetic group that includes the effects of perpendicular and parallel magnetic fluctuations. Comparing the results from the GFS and TGLF models, we find that GFS exhibits the best agreement with the database of CGYRO linear growth rates. Comparing the model results for the electromagnetic scans shows that GFS captures the effects of parallel magnetic fluctuations accurately, while the TGLF model does not, as it lacks sufficient perpendicular energy resolution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Implementation of Pfirsch–Schlüter parallel flow in x-ray imaging crystal spectrometer inversion analysis

The x-ray imaging crystal spectrometer (XICS) tomographic inversion code for Wendelstein 7-X (W7-X) has been modified to consider the effects of parallel flows and has been applied to analyze measurements taken during recent experimental campaigns. Previous analysis neglected the effects of parallel flows due to the primarily perpendicular geometry of the sightlines and the small magnitude predicted by neoclassical theory. To reconsider these effects, the incompressibility condition for plasma flows is used to calculate the parallel Pfirsch–Schlüter flow component for the equilibrium configuration. By incorporating this condition along with the geometry of the sightlines—i.e. the fractional contributions of perpendicular and parallel flows—, an updated expression for the measured flow is used for the profile inversion. Application of this modified inversion code to data from W7-X shows that the magnitude of the radial electric field and the flux surface averaged perpendicular flow are reduced by approximately a factor of 2 and brought into better agreement with neoclassical predictions and charge exchange recombination spectroscopy measurements.

Pfirsch–Schlüter flows↗

Effect of parallel flow on resonant layer responses in high beta plasmas

Abstract Resonant layers in a tokamak respond to non-axisymmetric magnetic perturbations by amplifying the mode amplitude and balancing the plasma rotation through magnetic reconnection and force balance, respectively. This resonant response can be characterized by local layer parameters and especially by a single quantity in the linear regime, the so-called inner-layer Δ. The computation of Δ under two-fluid drift-MHD formalism has been progressed by reducing the order of the system in the phase space, where the shielding current is approximated as being only carried by electrons, a posteriori . In this study, we relax the approximation and compute Δ accounted for by the parallel flow associated with the ion shielding current. The posteriori is numerically verified in great agreement with the original SLAYER developed in a previous paper (J.-K. Park 2022 Phys. Plasmas 29 072506). Extending the resonant layer response theory to high β plasmas, our research findings answer two important questions: how the parallel flow influences the resonant layer response and why the parallel flow effect appears in high β plasmas. The complicated plasma compression in high β regime allows the parallel flow response to give rise to the ion shielding current, which not only shifts the zero-crossing condition of the ExB flow but also enhances the field penetration threshold. Technically, the Riccati matrix transformation method is adapted to handle the numerical stiffness due to the increased order of the system. The high fidelity of this numerical method makes use of further extension of the model to higher-order systems to take other physical phenomena into account. This work is envisaged to predict the resonant layer response under high β fusion reactor conditions.

Lee, Yeongsun (ORCID:000000034474416X)↗

Scalable Parallel Measurement of Individual Nitrogen-Vacancy Centers

The nitrogen-vacancy (NV) center in diamond is a solid-state spin defect that has been widely adopted for quantum sensing and quantum information processing applications. Typically, experiments are performed either with a single isolated NV center or with an unresolved ensemble of many NV centers, resulting in a trade-off between measurement speed and spatial resolution or control over individual defects. In this work, we introduce an experimental platform that bypasses this trade-off by addressing multiple optically resolved NV centers in parallel. We perform charge- and spin-state manipulations selectively on multiple NV centers from within a larger set, and we manipulate and measure the electronic spin states of over 100 NV centers in parallel. We show that the high signal-to-noise ratio of the measurements enables the detection of shot-to-shot pairwise correlations between the spin states of 108 NV centers, corresponding to the simultaneous measurement of 5778 unique correlation coefficients. We discuss how our platform can be scaled to parallel experiments with thousands of individually resolved NV centers. These results enable parallelized high-throughput sensing experiments that retain the spatial resolution of single defects and will, thereby, help to unlock advances in applications such as single-molecule NMR and characterization of integrated circuits. In addition, our approach to multiplexing provides a natural platform for the application of recently developed correlated sensing techniques.

NV centers↗

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S↗

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan↗

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,↗

Steering law for parallel mounted double-gimbaled control moment gyros. Revision A

Mounting of double-gimbaled control moment gyros (CMG's) of unlimited outer gimbal angle freedom with all their outer gimbal axes parallel allows drastic simplification of the CMG steering law development hardware. The advantages of the parallel mounting for the CMG steering law development are such that a law could be developed which is applicable to any number of CMG's with arbitrary angular momentum. Parallel mounting of the CMG's in conjunction with the steering law can therefore be considered a CMG kit suitable for many missions of differing momentum requirements. It also means that increasing momentum demands during the design phase of a space vehicle can be easily met by the addition of one or more CMG's of the original momentum capacity rather than a redesign to a larger momentum capacity. Another advantage of the parallel mounting is that the failure of any CMG can be treated like any other, i.e., only one failure mode is possible. The CMG steering law distributes the CMG momentum vectors such that all inner gimbal angles are equal which reduces the rate requirements on the outer gimbal axes. The steering law also spreads the outer gimbals which ensures avoidance of singularities internal to the angular momentum envelope.

Kennel, H. F.↗

Iterative algorithms for large sparse linear systems on parallel computers

Algorithms for assembling in parallel the sparse system of linear equations that result from finite difference or finite element discretizations of elliptic partial differential equations, such as those that arise in structural engineering are developed. Parallel linear stationary iterative algorithms and parallel preconditioned conjugate gradient algorithms are developed for solving these systems. In addition, a model for comparing parallel algorithms on array architectures is developed and results of this model for the algorithms are given.

Adams, L. M.↗

Massively parallel processor computer

An apparatus for processing multidimensional data with strong spatial characteristics, such as raw image data, characterized by a large number of parallel data streams in an ordered array is described. It comprises a large number (e.g., 16,384 in a 128 x 128 array) of parallel processing elements operating simultaneously and independently on single bit slices of a corresponding array of incoming data streams under control of a single set of instructions. Each of the processing elements comprises a bidirectional data bus in communication with a register for storing single bit slices together with a random access memory unit and associated circuitry, including a binary counter/shift register device, for performing logical and arithmetical computations on the bit slices, and an I/O unit for interfacing the bidirectional data bus with the data stream source. The massively parallel processor architecture enables very high speed processing of large amounts of ordered parallel data, including spatial translation by shifting or sliding of bits vertically or horizontally to neighboring processing elements.

Fung, L. W.↗

Parallel, Asynchronous Executive (PAX): System concepts, facilities, and architecture

The Parallel, Asynchronous Executive (PAX) is a software operating system simulation that allows many computers to work on a single problem at the same time. PAX is currently implemented on a UNIVAC 1100/42 computer system. Independent UNIVAC runstreams are used to simulate independent computers. Data are shared among independent UNIVAC runstreams through shared mass-storage files. PAX has achieved the following: (1) applied several computing processes simultaneously to a single, logically unified problem; (2) resolved most parallel processor conflicts by careful work assignment; (3) resolved by means of worker requests to PAX all conflicts not resolved by work assignment; (4) provided fault isolation and recovery mechanisms to meet the problems of an actual parallel, asynchronous processing machine. Additionally, one real-life problem has been constructed for the PAX environment. This is CASPER, a collection of aerodynamic and structural dynamic problem simulation routines. CASPER is not discussed in this report except to provide examples of parallel-processing techniques.

Jones, W. H.↗

Multigrid solvers on parallel computers

Massively parallel computers, as considered in this investigation, are not yet available. However, a large-scale parallel computer cannot usefully be designed before the hypothetical algorithms which will employ it are studied. Most of the studies of parallel partial differential equations (PDE) solvers are based on solution techniques much slower (on sequential machines) than multigrid methods. Multigrid methods are highly parallelizable. Each of their processes can simultaneously be performed at all grid points. The present investigation is concerned with a preliminary exploration of the potential of multigrid, or, more generally, Multi-Level Adaptive Techniques (MLAT) on computers with many processors. Basic processes are considered, taking into account coarse-grid approximation, relaxation, coarse-grid corrections, full multigrid algorithms, nonlinear problems and eigenvalue problems, fine-to-coarse correction, and chains of problems. Details of parallel multigrid processing are also examined.

Brandt, A.↗

Finite element computation with parallel VLSI

This paper describes a parallel processing computer consisting of a 16-bit microcomputer as a master processor which controls and coordinates the activities of 8086/8087 VLSI chip set slave processors working in parallel. The hardware is inexpensive and can be flexibly configured and programmed to perform various functions. This makes it a useful research tool for the development of, and experimentation with parallel mathematical algorithms. Application of the hardware to computational tasks involved in the finite element analysis method is demonstrated by the generation and assembly of beam finite element stiffness matrices. A number of possible schemes for the implementation of N-elements on N- or n-processors (N is greater than n) are described, and the speedup factors of their time consumption are determined as a function of the number of available parallel processors.

Mcgregor, J.↗