Search NASASearch

SEARCH · Search NASA

Results for “Parallel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Radiological Source Term Estimation and Isotopic Identification with Parallel Log Domain Particle Filters

This paper presents a parallel log-domain particle filtering algorithm combined with gamma spectrum unfolding to perform localization, identification, and evaluation of multiple point sources of various isotopes in an environment with attenuating obstacles. The method uses sets of precomputed attenuation kernels that map the attenuation characteristics of the environment. These kernels are specific to the energy level of a photopeak of interest. The spectral measurements are deconvolved into count measurements of each photopeak. These count measurements are fed into a set of parallel particle filters using attenuation kernels computed for that photopeak’s energy level. The individual regularized particle filters perform all likelihood calculations in the logarithmic domain to mitigate the effects of particle degeneracy. The output of each particle filter is combined to estimate which isotopes are present as well as their positions and strengths. The performance of the algorithm is characterized in a lab-scale environment using a mobile robot equipped with a gamma ray spectrometer in the presence of up to three different radioactive isotopes simultaneously. The sources were localized to within 10 cm, and their strengths were estimated within 10% of their true values. Furthermore, the isotopes were all correctly identified, and no spurious sources were reported.

42 ENGINEERING

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)

Parallel Variable Population Multi-Objective Optimizer (pvpmoo) v1.0

This is a parallel variable population multi-objective optimizer with an adaptive unified differential evolution algorithm or a genetic algorithm. It can also be used for single objective optimization. Some features of this code include: 1) The population size varies from generation to generation to save the total # of objective function evaluations. 2) The population is uniformly distributed to a number of parallel processors for simultaneous objective function evaluation. 3) The objective function evaluation can be attained from an external simulation program with control variables in its input file and objectives calculated from its output files. 4) The optimizer includes an adaptive unified differential evolution algorithm and a real value genetic algorithm. The parameters in the unified differential evolution algorithm can be chosen to attain any mutation schemes in the published literature.

Qiang, Ji

Matrix-based Parallel Redistribution

MatRed is a parallel redistribution tool for HPC applications. It provides a simple approach that only requires a few relation matrices between entities to build redistribution matrices in parallel simulation codes. In particular, MatRed is well-suited for simulation codes based on finite element/volume methods.

Kalchev, DelyanZ [Lawrence Livermore National Labo

Design and Analysis of 10” Parallel Plate Relief Device

All Cryomodule (CM) and Cryogenic Distribution System (CDS) relieving into Helium Low Pressure (LP) return header, which is connected to compressor suction so, helium can be preserved during small flow relieving event and recirculated to system. However, during worst case scenario, Helium LP header requires a parallel plate relief device to relieve excess pressure from header. To complete the CDS Warm piping header, a new design for a 10 parallel plate relief device is necessary to relieve outside of the tunnel into atmosphere.

Chicas, Kelly

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING

Portable Parallel Algorithms and Frameworks for Exascale Graph Analytics

Graphs (or networks) are a tool used to model the interactions among various entities. Efficiently processing large graphs has recently attracted significant attention due to the applications of graphs in various domains, such as biology, chemistry, and cyber-security. Analyzing the structure and properties of these graphs is an important component of many scientific computing pipelines. With the explosion in the volume of data, graphs have become very large and can contain hundreds of billions of vertices and trillions of edges. Therefore, it is crucial to develop high-performance methods to enable graph analysis to be done quickly and energy-efficiently. Furthermore, these solutions should be highly parallel in order to take advantage of modern parallel machines. However, designing efficient solutions is not enough. With the wide variety of computing environments available, each with different programmability and performance characteristics, it is necessary to develop solutions that are portable in terms of both performance (i.e., provide theoretical guarantees) and programmability (i.e., provide high level abstractions).

97 MATHEMATICS AND COMPUTING

Oblique instability of quasi-parallel whistler waves in the presence of cold and warm electron populations

Whistler waves propagating nearly parallel to the ambient magnetic field experience a nonlinear instability due to transverse currents when the background plasma has a population of sufficiently low energy electrons. Intriguingly, this nonlinear process may generate oblique electrostatic waves, including whistlers near the resonance cone with properties resembling oblique chorus waves in the Earth’s magnetosphere. Focusing on the generation of oblique whistlers, earlier analysis of the instability is extended here to the case where low-energy background plasma consists of both a “cold” population with energy of a few eV and a “warm” electron component with energy of the order of 100 eV. This is motivated by spacecraft observations in the Earth’s magnetosphere where oblique chorus waves were shown to interact resonantly with the warm electrons. The main new results are: 1) the instability producing oblique electrostatic waves is sensitive to the shape of the electron distribution at low energies. In the whistler range of frequencies, two distinct peaks in the growth rate are typically present for the model considered: a peak associated with the warm electron population at relatively low wavenumbers and a peak associated with the cold electron population at relatively high wavenumbers; 2) overall, the instability producing oblique whistler waves near the resonance cone persists (with a reduced growth rate) even in the cases where the temperature of the cold population is relatively high, including cases where cold population is absent and only the warm population is included; 3) particle-in-cell simulations show that the instability leads to heating of the background plasma and formation of characteristic plateau and beam features in the parallel electron distribution function in the range of energies resonant with the instability. The plateau/beam features have been previously detected in spacecraft observations of oblique chorus waves. However, they have been attributed to external sources and have been proposed to be the mechanism generating oblique chorus. In the present scenario, the causality link is reversed and the instability generating oblique whistler waves is shown to be a possible mechanism for formation of the plateau and beam features.

79 ASTRONOMY AND ASTROPHYSICS

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)

Electron Influence on the Parallel Proton Firehose Instability in 10-moment, Multifluid Simulations

Instabilities driven by pressure anisotropy play a critical role in modulating the energy transfer in space and astrophysical plasmas. For the first time, we simulate the evolution and saturation of the parallel proton firehose instability using a multifluid model without adding artificial viscosity. These simulations are performed using a 10-moment, multifluid model with local and gradient relaxation heat-flux closures in high-β proton–electron plasmas. When these higher-order moments are included and pressure anisotropy is permitted to develop in all species, we find that the electrons have a significant impact on the saturation of the parallel proton firehose instability, modulating the proton pressure anisotropy as the instability saturates. Even for lower β's more relevant to heliospheric plasmas, we observe a pronounced electron energization in simulations using the gradient relaxation closure. Our results indicate that resolving the electron pressure anisotropy is important to correctly describe the behavior of multispecies plasma systems.

79 ASTRONOMY AND ASTROPHYSICS

Plastic Parallel Pathways Platform - 4P Model

Global momentum is building towards a circular economy capable of keeping plastics in use and out of waste streams. Given that 79% of all plastic produced since 1950 has accumulated in landfills or the natural environment,rapid implementation of various end-of-life (EoL) management technologies will be needed to reach this target. However, it can be challenging to develop an effective plastic EoL strategy when the available options - chemical or molecular recycling, energy recovery, upcycling, downcycling, closed-loop (plastic-to-plastic) or open-loop (plastic-to-x) recycling, among others - can generate products ranging from low-grade to virgin-quality plastic and from fuels to value-added chemicals. We present a flexible material flow model capable of analyzing the effects of both plastic-to-plastic and plastic-to-x EoL management strategies on the U.S. PET economy. This Plastic Parallel Pathways Platform (4P) assesses the environmental impacts, costs, and circularity of a PET system in which waste is managed through six potential EoL pathways: landfill, incineration with energy recovery, pyrolysis to fuel oil, upcycling to glass fiber reinforced plastic (GFRP), mechanical recycling to low-grade PET, and chemical recycling (glycolysis) to bottle-grade PET. We compare the pathways across multiple metrics using multi-criteria decision analysis (MCDA) and then use a brute force algorithm to predict an optimal combination of EoL pathways to minimize greenhouse gas (GHG) emissions and costs and maximize circularity. This work highlights the need to implement a diverse portfolio of EoL strategies in parallel to enable a PET economy that meets environmental, economic, and circularity requirements simultaneously.

downcycling

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING

Parallel computing for power system climate resiliency: Solving a large-scale stochastic capacity expansion problem with mpi-sppy

Here we propose a nodal stochastic generation and transmission expansion planning model that incorporates the output from high-resolution global climate models through load and generation availability scenarios. We implement our model in Pyomo and perform computational studies on a realistically-sized test case of the California electric grid in a high performance computing environment. We propose model reformulations and algorithm tuning to efficiently solve this large problem using a variant of the Progressive Hedging Algorithm. We utilize the parallelization capabilities and overall versatility of mpi-sppy, exploiting its hub-and-spoke architecture to concurrently obtain inner and outer bounds on an optimal expansion plan. Initial results show that instances with 360 representative days on a system with over 8,000 buses can be solved to within 5% of optimality in under 4 h of wall clock time, a first step towards solving a large-scale power system expansion planning problem across a wide range of climate-informed operational scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Rate-induced aging effects on Parallel-Plate Avalanche Counter (PPAC) caused by heavy ion beams

The Facility for Rare Isotope Beams (FRIB) is one of the premier scientific user facilities for nuclear science with radioactive beams, capable of producing most (approximately 80%) of the isotopes expected to exist, from oxygen to uranium, at energies up to 200 MeV/u. With the increase in beam power from the present 10 kW to the planned 400 kW, FRIB experiments are about to enter a new era. An unprecedented rate capability as well as stable performance of all the planned instrumentation intended for beam diagnostics and beam tuning is required at the expected high beam intensities (> 1 MHz). A summary of aging phenomena at high heavy-ion beam rates observed in the Advanced Rare Isotope Separator (ARIS) detectors for beam diagnostics, including Parallel Plate Avalanche Counters (PPAC) and plastic scintillation for time-of-flight measurements, is discussed. Current research and development project to mitigate rate-induced aging are presented.

Aging effects

Complex Fluid‐Driven Fractures Caused by Crack‐Parallel Stress

Abstract Managing fluid‐driven fracture networks is crucial for subsurface resource utilization, yet the current understanding of the key controlling factors remains insufficient. While geologic discontinuities have been shown to significantly influence fracture network complexity, this study identifies another major contributor. We conducted a new set of experiments using a transparent true triaxial cell, which enabled video recording of the temporal evolution of fluid‐driven fracture paths. Using pseudo‐2D samples without macroscale structural discontinuities, we observed multiple occurrences of hydraulic fracture curving and branching under anisotropic boundary stresses. We proposed a theoretical model demonstrating that the stress parallel to the crack line in the solid matrix near the crack tip (i.e., the T ‐stress) accounts for the observed fracture curving behavior. This finding suggests that T ‐stress is an additional mechanism contributing to the complexity of fluid‐driven fracture networks in the subsurface, besides the geologic discontinuities.

58 GEOSCIENCES

The impact of non-local parallel electron transport on plasma-impurity reaction rates in tokamak scrape-off layer plasmas

Abstract Plasma-impurity reaction rates are a crucial part of modelling tokamak scrape-off layer (SOL) plasmas. To avoid calculating the full set of rates for the large number of important processes involved, a set of effective rates are typically derived which assume Maxwellian electrons. However, non-local parallel electron transport may result in non-Maxwellian electrons, particularly close to divertor targets. Here, the validity of using Maxwellian-averaged rates in this context is investigated by computing the full set of rate equations for a fixed plasma background from kinetic and fluid SOL simulations. We consider the effect of the electron distribution as well as the impact of the electron transport model on plasma profiles. Results are presented for lithium, beryllium, carbon, nitrogen, neon and argon. It is found that electron distributions with enhanced high-energy tails can result in significant modifications to the ionisation balance and radiative power loss rates from excitation, on the order of 50%–75% for the latter. Fluid electron models with Spitzer-Härm or flux-limited Spitzer-Härm thermal conductivity, combined with Maxwellian electrons for rate calculations, can increase or decrease this error, depending on the impurity species and plasma conditions. Based on these results, we also discuss some approaches to experimentally observing non-local electron transport in SOL plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY