Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Process–Property–Performance Mapping of Additively Manufactured 316H Stainless Steel Components

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development of advanced materials and components fabricated via additive manufacturing, and is using laser powder bed fusion (LPBF) of 316H stainless steel as an initial case study. In the previous fiscal year, miniature high-throughput specimens were printed on multiple LPBF systems to provide initial processing windows to minimize porosity and limit epitaxial grain growth during prints. This fiscal year, scaled builds were completed on three different LPBF systems at ORNL: a GE Concept Laser M2, a Renishaw AM400, and an EOS M290. Builds on the Concept Laser were conducted on multiple powder lots and processing parameter ranges to provide microstructure effects on time-independent and time-dependent mechanical properties. Builds on the Renishaw were produced using Oak Ridge National Laboratory (ORNL)-optimized printing parameters and Argonne National Laboratory (ANL)-optimized printing parameters to compare outcomes of parallel process optimization efforts at different national laboratories on the same LPBF system. Similarly, the build completed on the EOS M290 replicated the processing parameters of builds completed at Los Alamos National Laboratory (LANL). Optical microscopy and electron backscatter diffraction characterization was completed on all builds. In addition to the general round robin characterization, this work-package generated time-independent data, including tensile and fracture toughness test data on scaled Concept Laser builds as a function of processing parameters and post-build heat treatment. This analysis is complimentary to work in parallel work packages aiming to establish heat treatment and processing effects on time-dependent properties. It was found that although the stress-relief heat treatment provides the highest strength at lower-temperatures, tensile strength begins to converge at higher temperatures regardless of heat treatment condition. In addition, the more rigorous solution annealing and hot-isostatic pressing post-build heat treatments result in higher fracture toughness than the stress-relieved condition. The root-causes of the lower fracture toughness of the stress-relieved LPBF 316H material was informed via a stress-relief optimization study on a scaled concept laser print, where it was found that although dislocation recovery was largely complete after only a couple hours at 650°C, the extended hold of the current 24h heat treatment employed on scaled builds likely caused increased carbide volume fractions along the LPBF 316H grain boundaries, thereby deteriorating crack propagation resistance. This trend was seen to become more deleterious with additional increases of stress-relief temperature to 750°C or 850°C. These results have helped inform a new optimal stress-relief annealing condition for LPBF 316H for future campaign testing (650°C for 2h).

36 MATERIALS SCIENCE↗

The Viskores User's Guide, Release 1.1

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created Viskores: the visualization toolkit for multi/many-core architectures. Viskores supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. Viskores also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although Viskores provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗

Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation

This Final Scientific and Technical Report summarizes work performed under the Phase IIA SBIR project “Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation” (DE-SC0020906) from August 2023 through August 2025. The objective of the Phase IIA effort was to mature and harden capabilities developed during Phase II, with the goal of enabling practical interoperability between Coreform’s isogeometric analysis (IGA) technologies and the Multiphysics Object-Oriented Simulation Environment (MOOSE), while improving robustness, performance, and scalability for complex, nuclear-relevant geometries. Over the course of Phase IIA, the project established and validated an extraction-based interoperability pathway between Coreform tools and MOOSE. A combined mesh and matrix format was defined collaboratively with MOOSE developers and integrated into the solver, enabling standard MOOSE workflows to operate on data exported from Coreform’s IGA and Flex Representation Method (FRM) pipelines. Early demonstrations validated architectural compatibility using linear solid mechanics problems, while later efforts focused on benchmark testing and external use. By the end of the project period, engineers at BWXT were able to independently set up and execute a simulation using the Coreform–MOOSE workflow and provide direct feedback that informed further refinement. In parallel, substantial effort was devoted to improving the robustness of trimmed U-spline construction for complex CAD geometries. A growing test suite of nuclear-relevant models was compiled through collaboration with multiple stakeholders and used to drive extensive bug fixing and reliability improvements. These efforts resulted in improved robustness and performance, including the addition of fallback capabilities that enhance reliability when the underlying commercial CAD kernel fails. Performance-oriented work progressed later in the project, with the development and demonstration of methods to decompose complex geometries into structured subregions and updated data representations to support more efficient solver processing. Additionally, extensive enhancements to threadsafe parallel data structures and trimming operations established a foundation for scalable processing of large assemblies. Collaboration with Sandia National Laboratories on the SGM geometric modeling kernel advanced to a functioning interface test case, positioning the workflow for future kernel integration. Overall, the Phase IIA effort successfully transitioned the project from architectural proof-of-concept to externally exercised, solver-integrated capability, while clarifying remaining technical challenges related to standardization, performance optimization, and kernel integration.

42 ENGINEERING↗

Exotic Uses of Neutrons an X-rays as Probes for Chiral Magnets (Early Career Award) (Final Report)

This is the final report for Exotic Uses of Neutrons an X-rays as Probes for Chiral Magnets, an Early Career award to Prof. Dustin Gilbert, University of Tennessee, running 09/01/2020 - 08/31/2025. This project used neutron scattering as a unique tool to investigate magnetic chiral structures. Neutron scattering is a powerful technique in which the neutron wavepacket scatters from nuclear or magnetic structures, providing insight into the structure of a material and its magnetic features. In this work, we focused on chiral magnetic structures. Most magnetic materials are colinear ferro- or antiferromagnets, where the spin moments align parallel or anti-parallel with their neighbors. In some materials, the magnetic moments instead curl and can form closed loops. These structures, with the additional feature that the core and perimeter are oriented in opposite out-of-plane directions, exhibit a property called topology; these structures are known as skyrmions. Topology is a broadly used term, but here it refers to a magnetic configuration that cannot be created or destroyed through any continuous transformation. For these looped structures, the closed loop cannot be destroyed continuously, giving rise to unique properties such as collective dynamics and particle-like behavior. This also raises fundamental questions about how such structures form and evolve.

36 MATERIALS SCIENCE↗

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

Mechanisms of regulation of the rhizosphere, roots and shoots of naive poplars

Trees are associated with a broad range of microorganisms colonising the diverse tissues of their host. However, the early dynamics of the microbiota assembly microbiota from the root to shoot axis and how it is linked to root exudates and metabolite contents of tissues remain unclear. Here, we characterised how fungal and bacterial communities are altering root exudates as well as root and shoot metabolomes in parallel with their establishment in poplar cuttings (Populus tremula x tremuloides clone T89) over 30 days of growth. Sterile poplar cuttings were planted in natural or gamma irradiated soils. Bulk and rhizospheric soils, root and shoot tissues were collected from day 1 to day 30 to track the dynamic changes of fungal and bacterial communities in the different habitats by DNA metabarcoding. Root exudates and root and shoot metabolites were analysed in parallel by gas chromatography-mass spectrometry.

09 BIOMASS FUELS↗

Angle-resolved polarized Raman spectroscopy study of phosphorene nanoribbons

We present a systematic angle-resolved polarized Raman spectroscopy (ARPRS) study of black phosphorus (BP) nanostructures formed via electrochemical sodium- and lithium-ion intercalation. Sodium intercalation leads to bundles of densely packed, highly uniform phosphorene nanoribbons (PNRs) separated by parallel amorphous channels, whereas lithium intercalation results in shorter, irregular nanoribbon-like segments with lower aspect ratios. In both cases, six additional Raman peaks (P1–P6) appear alongside the three primary Raman-active modes of BP (A$^{1}_{g}$, B 2g , and A$^{2}_{g}$). These peaks are attributed to the amorphous regions, as confirmed by their isotropic angular dependence in ARPRS measurements. The three BP modes show pronounced angular variations that differ significantly between the two intercalated samples. In sodium-intercalated BP, A$^{1}_{g}$ and A$^{2}_{g}$ modes retain a dumbbell-like angular dependence under parallel polarization with enhanced anisotropy and reduced symmetry under crossed polarization. At the same time, the B 2g mode transitions from four-lobed (cloverleaf) polar plot to a butterfly-like one. In contrast, lithium-intercalated BP exhibits weaker anisotropy and less distinct angular polar plots for all three modes. These differences reflect the sensitivity of phonon behavior to underlying nanostructure morphology. The vibrational frequencies density of states (FDOS) calculations attribute the B 2g mode transformation to phonon band folding and mode mixing in PNRs. This study demonstrates the power of ARPRS in probing phonon-structure relationships and highlights the influence of edge geometry and quantum confinement on phonon dispersion in PNRs.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

High y + Shear-Stress Turbulence Implementation for High Flux Isotope Reactor Narrow Channel Flows

The research objective of this work was to improve the engineering predictions of the turbulence characteristics of flows in curved narrow channels. Such channel flows are commonly encountered in nuclear research and test reactors, with one of them being the high-flux isotope reactor (HFIR). Research reactors bear high heat fluxes, and the proper computing of turbulence is paramount for safe and reliable reactor operation. The study builds on the results of a previous direct numerical simulation of turbulence to inform a well-known Reynolds-averaged Navier–Stokes shear-stress turbulence model and improves its accuracy in simulating parallel channel flows. A new formulation of the loss term in the dissipation conservation equation is suggested. Combined with high wall distance computational grids, the new implementation provides a fast-running flow solution, suitable for engineering purposes. Model generalization for parallel channel flows, in a broader range of frictional Reynolds numbers, is suggested by introducing a new form of the model constants.

CFD↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Electron Heating in the Transrelativistic Perpendicular Shocks of Tilted Accretion Flows

Abstract General relativistic magnetohydrodynamic (GRMHD) simulations of black hole tilted disks—where the angular momentum of the accretion flow at large distances is misaligned with respect to the black hole spin—commonly display standing shocks within a few to tens of gravitational radii from the black hole. In GRMHD simulations of geometrically thick, optically thin accretion flows, applicable to low-luminosity sources like Sgr A* and M87*, the shocks have transrelativistic speed, moderate plasma beta (the ratio of ion thermal pressure to magnetic pressure is β pi1 ∼ 1–8), and low sonic Mach number (the ratio of shock speed to sound speed is M s ∼ 1–6). We study such shocks with 2D particle-in-cell simulations, and we quantify the efficiency and mechanisms of electron heating for the special case of preshock magnetic fields perpendicular to the shock direction of propagation. We find that the postshock electron temperature T e2 exceeds the adiabatic expectation T e2,ad by an amount T e 2 / T e 2 , ad − 1 ≃ 0.0016 M s 3.6 , nearly independent of the plasma beta and of the preshock electron-to-ion temperature ratio T e1 / T i1 , which we vary from 0.1 to unity. We investigate the heating physics for M s ∼ 5–6 and find that electron superadiabatic heating is governed by magnetic pumping at T e1 / T i1 = 1, whereas heating by B -parallel electric fields (i.e., parallel to the local magnetic field) dominates at T e1 / T i1 = 0.1. Our results provide physically motivated subgrid prescriptions for electron heating at the collisionless shocks seen in GRMHD simulations of black hole accretion flows.

Astronomy & Astrophysics↗

Cholla-MHD: An Exascale-capable Magnetohydrodynamic Extension to the Cholla Astrophysical Simulation Code

Abstract We present an extension of the massively parallel, GPU native, astrophysical hydrodynamics code Cholla to magnetohydrodynamics (MHD). Cholla solves the ideal MHD equations in their Eulerian form on a static Cartesian mesh utilizing the Van Leer + constrained transport integrator, the HLLD Riemann solver, and reconstruction methods at second and third order. Cholla’s MHD module can perform ≈260 million cell updates per GPU-second on an NVIDIA A100 while using the HLLD Riemann solver and second order reconstruction. The inherently parallel nature of GPUs combined with increased memory in new hardware allows Cholla’s MHD module to perform simulations with resolutions ∼500 3 cells on a single high-end GPU (e.g., an NVIDIA A100 with 80 GB of memory). We employ GPU direct Message Passing Interface to attain excellent weak scaling on the exascale supercomputer Frontier, while using 74,088 GPUs and simulating a total grid size of over 7.2 trillion cells. A suite of test problems highlights the accuracy of Cholla’s MHD module and demonstrates that zero magnetic divergence in solutions is maintained to round off error. We also present new testing and CI tools using GoogleTest, GitHub Actions, and Jenkins that have made development more robust and accurate and ensure reliability in the future.

Astronomy & Astrophysics↗

A Model for Pair Production Limit Cycles in Pulsar Magnetospheres

Abstract It was recently proposed that the electric field oscillation as a result of self-consistent e ± pair production may be the source of coherent radio emission from pulsars. Direct particle-in-cell simulations of this process have shown that the screening of the parallel electric field by this pair cascade manifests as a limit cycle, as the parallel electric field is recurrently induced when pairs produced in the cascade escape from the gap region. In this work, we develop a simplified time-dependent kinetic model of e ± pair cascades in pulsar magnetospheres that can reproduce the limit-cycle behavior of pair production and electric field screening. This model includes the effects of a magnetospheric current, the escape of e ± , as well as the dynamic dependence of pair production rate on the plasma density and energy. Using this simple theoretical model, we show that the power spectrum of electric field oscillations averaged over many limit cycles is compatible with the observed pulsar radio spectrum.

Astronomy & Astrophysics↗

Fully Kinetic Simulations of Proton-beam-driven Instabilities from Parker Solar Probe Observations

The expanding solar wind plasma ubiquitously exhibits anisotropic nonthermal particle velocity distributions. Typically, proton velocity distribution functions (VDFs) show the presence of a core and a field-aligned beam. Novel observations made by the Parker Solar Probe (PSP) in the innermost heliosphere have revealed new complex features in the proton VDFs, namely anisotropic beams that sometimes experience perpendicular diffusion. In this study, we use a 2.5D fully kinetic simulation to investigate the stability of proton VDFs with anisotropic beams observed by PSP. Our setup consists of a core and an anisotropic beam population that drift with respect to each other. This configuration triggers a proton beam instability from which nearly parallel fast magnetosonic modes develop. Our results demonstrate that before this instability reaches saturation, the waves resonantly interact with the beam protons, causing perpendicular heating at the expense of the parallel temperature.

79 ASTRONOMY AND ASTROPHYSICS↗

Regulation of Solar Wind Electron Temperature Anisotropy by Collisions and Instabilities

Abstract Typical solar wind electrons are modeled as being composed of a dense but less energetic thermal “core” population plus a tenuous but energetic “halo” population with varying degrees of temperature anisotropies for both species. In this paper, we seek a fundamental explanation of how these solar wind core and halo electron temperature anisotropies are regulated by combined effects of collisions and instability excitations. The observed solar wind core/halo electron data in ( β ∥ , T ⊥ / T ∥ ) phase space show that their respective occurrence distributions are confined within an area enclosed by outer boundaries. Here, T ⊥ / T ∥ is the ratio of perpendicular and parallel temperatures and β ∥ is the ratio of parallel thermal energy to background magnetic field energy. While it is known that the boundary on the high- β ∥ side is constrained by the temperature anisotropy-driven plasma instability threshold conditions, the low- β ∥ boundary remains largely unexplained. The present paper provides a baseline explanation for the low- β ∥ boundary based upon the collisional relaxation process. By combining the instability and collisional dynamics it is shown that the observed distribution of the solar wind electrons in the ( β ∥ , T ⊥ / T ∥ ) phase space is adequately explained, both for the “core” and “halo” components.

Yoon, Peter H. (ORCID:0000000181343790)↗

Simultaneous Observation of Ion-scale Wave Packets with Opposite Polarizations and Their Implications on the Generation Region in the Inner Heliosphere

This paper reports a dispersion analysis of two wave packets simultaneously observed near the local proton gyrofrequency by the Parker Solar Probe. The observed wave event exhibits clear two-banded wave packets both propagating along the magnetic field, characterized by left-handed (L-mode) and right-handed (R-mode) polarizations simultaneously. By incorporating the Doppler shift effect into a linear dispersion analysis, we find two possible scenarios that explain these simultaneous opposite polarizations: (1) Two inherently L-mode waves in the plasma frame, propagate parallel and antiparallel to the solar wind velocity, with similar wave frequencies and wave numbers. The polarization of the antiparallel propagating wave reverses as it moves sunward in the plasma frame while still comoving with the solar wind in the stationary frame. This reversal manifests the polarization of the wave as an R-mode in the spacecraft frame. (2) Simultaneous L-mode and R-mode waves propagate parallel to the solar wind velocity, with different wave frequencies and wave numbers. Concurrent proton observations during the wave event reveal a dominant anisotropic ($T$⟂/$T$ ∥ > 1) core distribution with a drifting beam population. Estimation of the linear growth rate for both L-mode and R-mode waves suggests that both scenarios are plausible, indicating that the observation is near the wave-generation region. We explore the potential impact of these simultaneous waves on solar wind heating and scattering effects, hypothesizing that such waves might enhance efficiency compared to waves with a single wave packet, contingent upon the statistical significance of such waves.

79 ASTRONOMY AND ASTROPHYSICS↗

Temporal Properties of Compressible Magnetohydrodynamic Turbulence

Describing the temporal properties of compressible magnetohydrodynamic (MHD) turbulence is a fundamental problem that has important implications for particle acceleration and transport in astrophysical plasmas. Here, by carefully analyzing the spatial and temporal properties of compressible MHD turbulence, we derive a new spectral power density function that is supported by simulations. This new function reveals that the low-frequency fluctuations are dominated by modes with small parallel wavenumbers with respect to the mean background magnetic field. Furthermore, for fluctuations with dynamically significant parallel wavenumbers, broadening around their eigenfrequencies is described by this function, which is in close agreement with simulations. We use this formalism to present the scaling properties of individual MHD modes. Such broadening is a direct consequence of nonlinear processes and is different for the three fundamental MHD modes. Our results provide a new window to investigate the temporal properties of turbulence and will enable further studies on the interaction between compressible MHD turbulence and energetic plasmas.

79 ASTRONOMY AND ASTROPHYSICS↗

Planar Collisionless Shock Simulations with the Semi-implicit Particle-in-cell Model FLEKS

This study investigates the applicability of the semi-implicit particle-in-cell code FLexible Exascale Kinetic Simulator (FLEKS) to heliospheric shock simulations. We examine one- and two-dimensional local planar shock simulations, initialized using MHD states with upstream conditions representative of plasmas in the hypersonic, β ∼ 1 regime, for both quasi-perpendicular and quasi-parallel configurations. The refined algorithm in FLEKS proves robust, enabling accurate shock simulations with a grid resolution on the order of the electron inertial length d e . Our simulations successfully capture key shock features, including shock structures (foot, ramp, overshoot, and undershoot), upstream and downstream waves (fast magnetosonic, whistler, Alfvén ion-cyclotron, and mirror modes), and non-Maxwellian particle distributions. Crucially, we find that at least two spatial dimensions are critical for accurately reproducing downstream-wave physics in quasi-perpendicular shocks and capturing the complex dynamics of quasi-parallel shocks, including surface rippling, shocklets, short, large-amplitude magnetic structures, magnetic reconnection, and jets. Furthermore, our parameter studies demonstrate the impact of mass ratio and grid resolution on shock physics. This work provides valuable guidance for selecting appropriate physical and numerical parameters for shock simulations using a semi-implicit PIC method, paving the way for incorporating kinetic shock processes into large-scale collisionless plasma simulations with the MHD-AEPIC model.

plasma astrophysics↗