Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Pedestal origin and extrapolation of high-density small edge-localised-modes peak parallel energy fluence in ITER and SPARC

Experimental analysis and simulations with the BOUT++ code show that small edge-localised modes (ELMs) in reactor-relevant high-density regimes originate in a region close to the separatrix and only marginally perturb the pedestal structure. The measured divertor peak parallel energy fluence (ε ∥,peak ) for a database of small ELM scenarios in DIII-D and ASDEX Upgrade can be reproduced, within 40 % accuracy on average, if an ad hoc modification of the Eich peak parallel ELM energy fluence model is applied to account for the small ELM pedestal birth location. This allows for first-order extrapolation of small-ELM divertor ε ∥,peak to ITER and SPARC, resulting in values that satisfy the nominal melting threshold of tungsten monoblocks of 12 MJ m −2 . The findings reported in this study, both via modelling and direct measurements, constitute a step forward in assessing small ELMs in high edge-collisionality scenarios as a viable plasma regime for the operation of next-generation fusion machines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Cyclotron breaking: a mechanism for parallel ion cyclotron waves to heat the fast solar wind

The Parker Solar Probe mission has observed near-continuous power in parallel ion cyclotron waves (PICWs) in the young, fast solar wind. These waves are unlikely to be directly produced by the turbulent cascade and are likely born of a local instability; yet, they are observed to both cool – and heat – the plasma. We propose that these observations can be self-consistently explained as the natural consequence of PICWs propagating in the inhomogeneous solar wind after they have been driven unstable. In this work, we argue that strong proton heating by a turbulent cascade of oblique ICWs will result in PICWs being driven unstable in a process known as quasi-linear focusing. Because the power in the turbulent cascade is concentrated at scales above the turbulent transition region, PICWs will be driven unstable within a range of wavenumbers parallel to the background magnetic field, 𝑘 ∥ , that is bounded from above by 𝑘$^{∗}_{∥P}$, corresponding to the start of the transition region. As unstable PICWs propagate away from the Sun to regions of lower proton density, their 𝑘 ∥ , multiplied by the proton inertial length 𝑑 p , increases. Eventually, 𝑘$^{∗}_{∥P}$ of the PICWs becomes larger than 𝑘$^{∗}_{∥P}$⁢𝑑 p and the waves damp, heating the solar wind. We call this effect ‘cyclotron breaking’, in analogy with ocean waves breaking on the shore. We then discuss the testable predictions of the theory, including a distinct heating signature in which PICWs cool fast protons and heat slow protons at any given heliocentric distance 𝑟. Finally, we conjecture that cyclotron breaking can lead to net heating by PICWs if the power emitted as PICWs decreases sufficiently rapidly with 𝑟 that local emission of PICWs is overwhelmed by the local damping of PICWs generated closer to the Sun.

plasma heating↗

Substituent and Heteroatom Effects on π–π Interactions: Evidence That Parallel-Displaced π-Stacking is Not Driven by Quadrupolar Electrostatics

Stacking interactions are a recurring motif in supramolecular chemistry and biochemistry, where a persistent theme is a preference for parallel-displaced aromatic rings rather than face-to-face π-stacking. This is usually explained in terms of quadrupole–quadrupole interactions between the arene moieties but that interpretation is inconsistent with accurate calculations, which reveal that the quadrupolar picture is qualitatively wrong. At typical π-stacking distances, quadrupolar electrostatics may differ in sign from an exact calculation based on charge densities of the interacting arenes. We apply symmetry-adapted perturbation theory to dimers composed of substituted benzene and various aromatic heterocycles, which display a wide range of electrostatic interactions, and we investigate the interplay of Pauli repulsion, dispersion, and electrostatics as it pertains to parallel-displaced π-stacking. Profiles of energy components along cofacial slip-stacking coordinates support a prominent role for the “van der Waals model” (dispersion in competition with Pauli repulsion), even for polar monomers where electrostatic interactions are significant. While electrostatic interactions are necessary to explain the optimal face-to-face π-stacking distance and to account for the relative orientation of one polar arene with respect to another, we find no evidence to support continued invocation of quadrupolar electrostatics as a basis for π-stacking. Our results suggest that a driving force for offset-stacking exists even in the absence of electrostatic interactions. Consequently, tuning electrostatics via functionalization does not guarantee that slip-stacking can be avoided. This has implications for rational design of soft materials and other supramolecular architectures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Kelvin–Helmholtz instability under stabilizing parallel magnetic field in nonhomogeneous compressible MHD flows

We study the Kelvin–Helmholtz instability (KHI) for the general case of a compressible, nonhomogeneous, magnetized plasma flow. The study is limited to a vortex sheet interface with an imposed parallel magnetic field. We introduce a new formalism based on a convective Mach number M c , a convective Alfvénic Mach number M Ac , and a total convective Mach number that combines the two. We derive an analytic expression of the KHI growth rate for a homogeneous flow (i.e., zero Atwood number, A=0) that converges toward both the expression for unmagnetized compressible flow and Chandrasekhar's expression for magnetized incompressible flow. Otherwise, the dispersion relation is solved numerically and allows deriving general stability diagrams of magnetized KHI for the triplet (A, M c , β −plasma) parameters. We show these parameters uniquely define all configurations for a parallel magnetic field. We also construct diagrams with respect to the convective Alfvénic Mach number, the β − plasma parameter, or the magnetic field showing which magnetic field strength is required for stabilizing a given shear flow. The theoretical growth rates are compared with 18 simulations made with the GAMERA code, currently used for 3D magnetospheric simulations. Finally, we apply our results to the analysis of a past KHI experiment performed at the OMEGA laser facility, showing linear theory succeeds to provide accurate estimates of the growth rate at early times. We further discuss how our results can inform future experiments in the high-Mach magnetized regime at the National Ignition Facility. Possible limitations of the study due to resistive, mixing, or turbulence effects are discussed.

compressible flows↗

Semi-implicit continuum kinetic modeling of weakly collisional parallel transport in a magnetic mirror

We present implicit-explicit (IMEX) kinetic simulations of weakly collisional parallel plasma transport in magnetic mirror configurations using the continuum code COGENT. The numerical scheme employs a Jacobian-free Newton–Krylov method with algebraic multigrid preconditioning to overcome the severe time step limitations imposed by strong mirror forces in fully explicit schemes. Applied to parameters relevant to the Wisconsin HTS Axisymmetric Mirror experiment, the IMEX approach enables time steps up to 2.5×10 4 times larger than those permitted by explicit methods, resulting in a 2500× speedup in 1D–2V simulations of parallel transport with kinetic ions and Boltzmann electrons. Additionally, a reduced bounce-averaged model for a square mirror is implemented to support the computationally intensive fully kinetic simulations. The bounce-averaged formulation is used to evaluate the numerical convergence of the velocity-space discretization algorithms and to assess the role of the collision model by comparing simulations employing the nonlinear Fokker–Planck and the simplified Lenard–Bernstein–Dougherty collision operators.

Collision theories↗

Optimizing temperature distributions for training neural quantum states using parallel tempering

Parametrized artificial neural networks (ANNs) can be very expressive ansatzes for variational algorithms, reaching state-of-the-art energies on many quantum many-body Hamiltonians. Nevertheless, the training of the ANN can be slow and stymied by the presence of local minima in the parameter landscape. One approach to mitigate this issue is to use parallel tempering methods, and in this work, we focus on the role played by the temperature distribution of the parallel tempering replicas. Using an adaptive method that adjusts the temperatures in order to equate the exchange probability between neighboring replicas, we show that this temperature optimization can significantly increase the success rate of the variational algorithm with negligible computational cost by eliminating bottlenecks in the replicas' random walk. Furthermore, we demonstrate this using two different neural networks, a restricted Boltzmann machine and a feedforward network, which we use to study a toy problem based on a permutation invariant Hamiltonian with a pernicious local minimum and the 𝐽 1 −𝐽 2 model on a rectangular lattice.

Neural network simulations↗

Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study

Many parallel and distributed computing research results are obtained in simulation, using simulators that mimic real-world executions on some target system. Each such simulator is configured by picking values for parameters that define the behavior of the underlying simulation models it implements. The main concern for a simulator is accuracy: simulated behaviors should be as close as possible to those observed in the real-world target system. This requires that values for each of the simulator's parameters be carefully picked, or “calibrated,” based on ground-truth real-world executions. Examining the current state of the art shows that simulator calibration, at least in the field of parallel and distributed computing, is often undocumented (and thus perhaps often not performed) and, when documented, is described as a labor-intensive, manual process. In this work we evaluate the benefit of automating simulation calibration using simple algorithms. Specifically, we use a real-world case study from the field of High Energy Physics and compare automated calibration to calibration performed by a domain scientist. Our main finding is that automated calibration is on par with or significantly outperforms the calibration performed by the domain scientist. Furthermore, automated calibration makes it straightforward to operate desirable tradeoffs between simulation accuracy and simulation speed.

Mc donald, Jesse↗

Stability Analysis of Parallel Connected Bidirectional WPT System

This paper presents a stability analysis of parallel-connected bi-directional series-series resonant network wireless power transfer (WPT), optimized for Electric Vehicle (EV) charging and vehicle-to-grid (V2G) applications. The study addresses critical stability challenges in systems integrated with diverse distributed energy resources (DERs), including photovoltaics, fuel cells, wind turbines, energy storage systems, and the AC grid. The stability of such integrated DC grid systems is paramount for ensuring reliable operation, particularly under varying power flow conditions and dynamic interactions between parallel WPT systems. The analysis included system impedance characterization, state-space modeling, and open and closed-loop stability evaluations. The results demonstrated that the integration of a robust control architecture effectively mitigates instability risks and supports scalable, efficient operation. This work underscores the converter's adaptability and its potential for large-scale deployment in wireless EV charging infrastructures and integrated DC grid systems.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Radiological Source Term Estimation and Isotopic Identification with Parallel Log Domain Particle Filters

This paper presents a parallel log-domain particle filtering algorithm combined with gamma spectrum unfolding to perform localization, identification, and evaluation of multiple point sources of various isotopes in an environment with attenuating obstacles. The method uses sets of precomputed attenuation kernels that map the attenuation characteristics of the environment. These kernels are specific to the energy level of a photopeak of interest. The spectral measurements are deconvolved into count measurements of each photopeak. These count measurements are fed into a set of parallel particle filters using attenuation kernels computed for that photopeak’s energy level. The individual regularized particle filters perform all likelihood calculations in the logarithmic domain to mitigate the effects of particle degeneracy. The output of each particle filter is combined to estimate which isotopes are present as well as their positions and strengths. The performance of the algorithm is characterized in a lab-scale environment using a mobile robot equipped with a gamma ray spectrometer in the presence of up to three different radioactive isotopes simultaneously. The sources were localized to within 10 cm, and their strengths were estimated within 10% of their true values. Furthermore, the isotopes were all correctly identified, and no spurious sources were reported.

42 ENGINEERING↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Parallel Variable Population Multi-Objective Optimizer (pvpmoo) v1.0

This is a parallel variable population multi-objective optimizer with an adaptive unified differential evolution algorithm or a genetic algorithm. It can also be used for single objective optimization. Some features of this code include: 1) The population size varies from generation to generation to save the total # of objective function evaluations. 2) The population is uniformly distributed to a number of parallel processors for simultaneous objective function evaluation. 3) The objective function evaluation can be attained from an external simulation program with control variables in its input file and objectives calculated from its output files. 4) The optimizer includes an adaptive unified differential evolution algorithm and a real value genetic algorithm. The parameters in the unified differential evolution algorithm can be chosen to attain any mutation schemes in the published literature.

Qiang, Ji↗

Matrix-based Parallel Redistribution

MatRed is a parallel redistribution tool for HPC applications. It provides a simple approach that only requires a few relation matrices between entities to build redistribution matrices in parallel simulation codes. In particular, MatRed is well-suited for simulation codes based on finite element/volume methods.

Kalchev, DelyanZ [Lawrence Livermore National Labo↗

Design and Analysis of 10” Parallel Plate Relief Device

All Cryomodule (CM) and Cryogenic Distribution System (CDS) relieving into Helium Low Pressure (LP) return header, which is connected to compressor suction so, helium can be preserved during small flow relieving event and recirculated to system. However, during worst case scenario, Helium LP header requires a parallel plate relief device to relieve excess pressure from header. To complete the CDS Warm piping header, a new design for a 10 parallel plate relief device is necessary to relieve outside of the tunnel into atmosphere.

Chicas, Kelly↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

Portable Parallel Algorithms and Frameworks for Exascale Graph Analytics

Graphs (or networks) are a tool used to model the interactions among various entities. Efficiently processing large graphs has recently attracted significant attention due to the applications of graphs in various domains, such as biology, chemistry, and cyber-security. Analyzing the structure and properties of these graphs is an important component of many scientific computing pipelines. With the explosion in the volume of data, graphs have become very large and can contain hundreds of billions of vertices and trillions of edges. Therefore, it is crucial to develop high-performance methods to enable graph analysis to be done quickly and energy-efficiently. Furthermore, these solutions should be highly parallel in order to take advantage of modern parallel machines. However, designing efficient solutions is not enough. With the wide variety of computing environments available, each with different programmability and performance characteristics, it is necessary to develop solutions that are portable in terms of both performance (i.e., provide theoretical guarantees) and programmability (i.e., provide high level abstractions).

97 MATHEMATICS AND COMPUTING↗

Oblique instability of quasi-parallel whistler waves in the presence of cold and warm electron populations

Whistler waves propagating nearly parallel to the ambient magnetic field experience a nonlinear instability due to transverse currents when the background plasma has a population of sufficiently low energy electrons. Intriguingly, this nonlinear process may generate oblique electrostatic waves, including whistlers near the resonance cone with properties resembling oblique chorus waves in the Earth’s magnetosphere. Focusing on the generation of oblique whistlers, earlier analysis of the instability is extended here to the case where low-energy background plasma consists of both a “cold” population with energy of a few eV and a “warm” electron component with energy of the order of 100 eV. This is motivated by spacecraft observations in the Earth’s magnetosphere where oblique chorus waves were shown to interact resonantly with the warm electrons. The main new results are: 1) the instability producing oblique electrostatic waves is sensitive to the shape of the electron distribution at low energies. In the whistler range of frequencies, two distinct peaks in the growth rate are typically present for the model considered: a peak associated with the warm electron population at relatively low wavenumbers and a peak associated with the cold electron population at relatively high wavenumbers; 2) overall, the instability producing oblique whistler waves near the resonance cone persists (with a reduced growth rate) even in the cases where the temperature of the cold population is relatively high, including cases where cold population is absent and only the warm population is included; 3) particle-in-cell simulations show that the instability leads to heating of the background plasma and formation of characteristic plateau and beam features in the parallel electron distribution function in the range of energies resonant with the instability. The plateau/beam features have been previously detected in spacecraft observations of oblique chorus waves. However, they have been attributed to external sources and have been proposed to be the mechanism generating oblique chorus. In the present scenario, the causality link is reversed and the instability generating oblique whistler waves is shown to be a possible mechanism for formation of the plateau and beam features.

79 ASTRONOMY AND ASTROPHYSICS↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

Electron Influence on the Parallel Proton Firehose Instability in 10-moment, Multifluid Simulations

Instabilities driven by pressure anisotropy play a critical role in modulating the energy transfer in space and astrophysical plasmas. For the first time, we simulate the evolution and saturation of the parallel proton firehose instability using a multifluid model without adding artificial viscosity. These simulations are performed using a 10-moment, multifluid model with local and gradient relaxation heat-flux closures in high-β proton–electron plasmas. When these higher-order moments are included and pressure anisotropy is permitted to develop in all species, we find that the electrons have a significant impact on the saturation of the parallel proton firehose instability, modulating the proton pressure anisotropy as the instability saturates. Even for lower β's more relevant to heliospheric plasmas, we observe a pronounced electron energization in simulations using the gradient relaxation closure. Our results indicate that resolving the electron pressure anisotropy is important to correctly describe the behavior of multispecies plasma systems.

79 ASTRONOMY AND ASTROPHYSICS↗