Search NASASearch

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING

Computational flow modeling of triply periodic minimal surfaces as feed channel spacers in ultra-high pressure reverse osmosis applications

Triply periodic minimal surfaces (TPMS) are a special class of mathematical surfaces characterized by a high surface area-to-volume ratio. They have generated considerable interest in fields such as acoustics, heat transfer, and membrane-based filtration processes. This study evaluates the performance of four different TPMS designs—Schoen Gyroid, Schoen Crossed Layers of Parallels (CLP), Schoen Transverse Crossed Layers of Parallels (tCLP), and Schwarz-Primitive—when used as feed channel spacers under ultra-high pressure reverse osmosis (UHPRO) conditions, at approximately 200 bar. Our experimentally validated computational fluid dynamics model reveal different flow patterns within the feed channels for each of the four TPMS designs, leading to varying hydrodynamic and permeation properties. Under the simulated UHPRO conditions, the Gyroid and tCLP designs yield up to a 23% increase in average permeate velocity and a 14% reduction in average membrane-surface concentration relative to a non-woven spacer of the same porosity. Furthermore, the enhanced performance comes with an increased feed channel pressure drop, although it only constitutes less than 4% of the operating pressure when extrapolated for a meter-long membrane module. Additionally, the study analyzes the effects of varying inlet velocity and spacer porosity on membrane performance. Overall, this research provides valuable insights into the potential use of TPMS spacers in UHPRO applications.

36 MATERIALS SCIENCE

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management

Distributed Augmentation, Hypersweeps, and Branch Decomposition of Contour Trees for Scientific Exploration

Contour trees describe the topology of level sets in scalar fields and are widely used in topological data analysis and visualization. A main challenge of utilizing contour trees for large-scale scientific data is their computation at scale using highperformance computing. To address this challenge, recent work has introduced distributed hierarchical contour trees for distributed computation and storage of contour trees. However, effective use of these distributed structures in analysis and visualization requires subsequent computation of geometric properties and branch decomposition to support contour extraction and exploration. In this work, we introduce distributed algorithms for augmentation, hypersweeps, and branch decomposition that enable parallel computation of geometric properties, and support the use of distributed contour trees as query structures for scientific exploration. Finally, we evaluate the parallel performance of these algorithms and apply them to identify and extract important contours for scientific visualization.

97 MATHEMATICS AND COMPUTING

Angle-Resolved Polarized Raman Study of Layered Cr 2 Se 3

The polarization-resolved Raman spectra of two-dimensional Cr 2 Se 3 synthesized via chemical vapor deposition (CVD) and chemical vapor transport (CVT) techniques were investigated in detail. The samples were characterized using X-ray diffraction (XRD), transmission electron microscopy (TEM), and energy-dispersive X-ray spectroscopy (EDS). A distinct polarization dependence was observed in the Raman intensity of all the Cr-Cr, Cr-Se, and Se-Se modes in both samples. The observed angle-dependent Raman intensities of each peak could be related to the crystal structure-specific Raman tensor. XRD results of the bulk Cr 2 Se 3 sample synthesized via CVT confirm its trigonal crystal structure, and the Raman peaks can be fitted using the Raman tensors for the A g and E g modes for both the parallel and crossed polarizations. However, for the Cr 2 Se 3 samples directly grown on Si/SiO 2 substrates by CVD, it was necessary to assume the triclinic crystal structure in order to explain the polarized Raman dependence of all the peaks in both parallel and crossed polarization directions. Furthermore, this is the first experimental result suggesting the existence of triclinic Cr 2 Se 3 crystal structure, which has been theoretically predicted in the Materials Project database.

36 MATERIALS SCIENCE

A Scalable Interior‐Point Gauss–Newton Method for PDE‐Constrained Optimization With Bound Constraints

Here, we present a scalable approach to solve a class of partial differential equation (PDE)‐constrained optimization problems with bound constraints. This approach utilizes a robust full‐space interior‐point (IP)‐Gauss–Newton optimization method. To cope with the poorly‐conditioned IP‐Gauss–Newton saddle‐point linear systems that need to be solved approximately, once per optimization step, we propose two spectrally related preconditioners. These preconditioners leverage the limited informativeness of data in regularized PDE‐constrained optimization problems. A block Gauss–Seidel preconditioner is proposed for the GMRES‐based solution of the IP‐Gauss–Newton linear systems. It is shown, for a large‐class of PDE‐ and bound‐constrained optimization problems, that the spectrum of the block Gauss–Seidel preconditioned IP‐Gauss–Newton matrix is asymptotically independent of discretization and is not impacted by the ill‐conditioning that notoriously plagues interior‐point methods. We exploit symmetry of the IP‐Gauss–Newton linear systems and propose a regularization and log‐barrier Hessian preconditioner for the preconditioned conjugate gradient (PCG)‐based solution of the equivalent IP‐Gauss–Newton–Schur complement linear systems. The eigenvalues of the block Gauss–Seidel preconditioned IP‐Gauss–Newton matrix, that are not equal to one, are identical to the eigenvalues of the regularization and log‐barrier Hessian preconditioned Schur complement matrix. The scalability of the approach is demonstrated on two example problems. The numerical solution of these optimization problems is shown to require a discretization independent number of IP‐Gauss–Newton linear solves. Furthermore, the linear systems are solved in a discretization and IP ill‐conditioning independent number of preconditioned Krylov subspace iterations. The parallel scalability of the preconditioner, achieved via algebraic multigrid component solvers when applicable, and the aforementioned algorithmic scalability permits a parallel scalable means to compute solutions of a large class of PDE‐ and bound‐constrained problems.

PDE-constrained optimization

Fragme∩t: An Open‐Source Framework for Multiscale Quantum Chemistry Based on Fragmentation

Fragment-based quantum chemistry offers a means to circumvent the nonlinear computational scaling of conventional electronic structure calculations, by partitioning a large calculation into smaller subsystems then considering the many-body interactions between them. Variants of this approach have been used to parameterize classical force fields and machine learning potentials, applications that benefit from interoperability between quantum chemistry codes. However, there is a dearth of software that provides interoperability yet is purpose-built to handle the combinatorial complexity of fragment-based calculations. To fill this void we introduce “Fragme∩t”, an open-source software application that provides a tool for community validation of fragment-based methods, a platform for developing new approximations, and a framework for analyzing many-body interactions. Fragme∩t includes algorithms for automatic fragment generation and structure modification, and for distance- and energy-based screening of the requisite subsystems. Checkpointing, database management, and parallelization are handled internally and results are archived in a portable database. Interfaces to various quantum chemistry engines are easy to write and exist already for Q-Chem, PySCF, xTB, Orca, CP2K, MRCC, Psi4, NWChem, GAMESS, and MOPAC. Applications reported here demonstrate parallel efficiencies around 96% on more than 1000 processors but also showcase that the code can handle large-scale protein fragmentation using only workstation hardware, all with a codebase that is designed to be usable by non-experts. Fragme∩t conforms to modern software engineering best practices and is built upon well established technologies including Python, SQLite, and Ray. The source code is available under the Apache 2.0 license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

CI/CD Efforts for Validation, Verification and Benchmarking OpenMP Implementations

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.

Jarmusch, Aaron

A closer look in the mirror: reflections on the matter/dark matter coincidence

We argue that the striking similarity between the cosmic abundances of baryons and dark matter, despite their very different astrophysical behavior, strongly motivates the scenario in which dark matter resides within a rich dark sector parallel in structure to that of the standard model. The near cosmic coincidence is then explained by an approximate ℤ$_{2}$ exchange symmetry between the two sectors, where dark matter consists of stable dark neutrons, with matter and dark matter asymmetries arising via parallel WIMP baryogenesis mechanisms. Taking a top-down perspective, we point out that an adequate ℤ$_{2}$ symmetry necessitates solving the electroweak hierarchy problem in each sector, without our committing to a specific implementation. A higher-dimensional realization in the far UV is presented, in which the hierarchical couplings of the two sectors and the requisite ℤ$_{2}$-breaking structure arise naturally from extra-dimensional localization and gauge symmetries. We trace the cosmic history, paying attention to potential pitfalls not fully considered in previous literature. Residual ℤ$_{2}$-breaking can very plausibly give rise to the asymmetric reheating of the two sectors, needed to keep the cosmological abundance of relativistic dark particles below tight bounds. We show that, despite the need to keep inter-sector couplings highly suppressed after asymmetric reheating, there can naturally be order-one couplings mediated by TeV scale particles which can allow experimental probes of the dark sector at high energy colliders. Massive mediators can also induce dark matter direct detection signals, but likely at or below the neutrino floor.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Influence of build direction on the fracture mechanism of 3D printed octet lattices

Here, we investigate the effects of 3D printing build directions on the fracture properties of octet lattice metamaterials made of polylactic acid (PLA) and how these effects vary with the relative density of the lattices. Single-edge notch bend samples are 3D printed in two orthogonal build directions at various relative densities. Our results show that the work of fracture for octet lattices with build directions parallel to the crack plane is significantly higher than those with build directions perpendicular to the crack plane. We also observed that the ratio of specific work of fracture between the two build directions remains nearly constant across different relative densities. In contrast, the ratio of peak load between the two build directions decreases as relative density increases. Furthermore, the build direction dictates the fracture mechanism. While the perpendicular build direction predominantly results in a brittle (mode I) fracture, the parallel build direction leads to a complex, delamination-dominated (mode II type) fracture. This phenomenon is largely governed by the weak interfaces formed between the printed layers and their interaction with the lattice geometry. These results reveal that the build direction governs the fracture mechanism and work of fracture of these lattice metamaterials and is therefore an important design consideration.

3D printed PLA

A scalable framework for efficient coupling of thermal and microstructural simulations in additive manufacturing

Predicting microstructure evolution in metal additive manufacturing (AM) is important for process optimization, but spatiotemporal scale disparities between thermal transport and microstructure evolution create significant challenges for efficient data transfer between simulation codes. To address this, we present Stork, a scalable framework for coupling thermal and microstructural simulations. Stork uses a sparse data representation to identify and store active solidification sub-volumes, enabling highly parallel quad-linear interpolation from coarse thermal grids to fine microstructure grids without large intermediate storage. We demonstrate the framework by coupling the semi-analytic heat transfer code 3DThesis with the time-parallel cellular automata code Toucan. This approach achieves over two orders of magnitude reduction in data generation time and file size compared to prior workflows. Numerical studies show that quad-linear interpolation preserves grain morphology and crystallographic texture in laser powder bed fusion (LPBF) simulations for coarsening ratios up to 16. Overall, Stork provides a scalable pathway for high-throughput, component-scale AM simulations on modern high-performance computing systems.

36 MATERIALS SCIENCE

T RI M E ++: Multi-threaded triangular meshing in two dimensions

We present T RI M E ++, a multi-threaded software library designed for generating two-dimensional meshes for intricate geometric shapes using the Delaunay triangulation. Multi-threaded parallel computing is implemented throughout the meshing procedure, making it suitable for fast generation of large-scale meshes. Three iterative meshing algorithms are implemented: the DistMesh algorithm, the centroidal Voronoi diagram meshing, and a hybrid of the two. We compare the performance of the three meshing methods in T RI M E ++, and show that the hybrid method retains the advantages of the other two. The software library achieves significant parallel speedup when generating large-scale meshes containing between 10 4 to 10 7 points. T RI M E ++ can handle complicated geometries and generates adaptive meshes of high quality.

97 MATHEMATICS AND COMPUTING

A provably stable numerical method for the anisotropic diffusion equation in confined magnetic fields

We present a novel numerical method for solving the anisotropic diffusion equation in magnetic fields confined to a periodic box which is accurate and provably stable. We derive energy estimates of the solution of the continuous initial boundary value problem. A discrete formulation is presented using operator splitting in time with the summation by parts finite difference approximation of spatial derivatives for the perpendicular diffusion operator. Weak penalty procedures are derived for implementing both boundary conditions and parallel diffusion operator obtained by field line tracing. We prove that the fully-discrete approximation is unconditionally stable. Discrete energy estimates are shown to match the continuous energy estimate given the correct choice of penalty parameters. A nonlinear penalty parameter is shown to provide an effective method for tuning the parallel diffusion penalty and significantly minimises rounding errors. Several numerical experiments, using manufactured solutions, the “NIMROD benchmark” problem and a single island problem, are presented to verify numerical accuracy, convergence, and asymptotic preserving properties of the method. Finally, we present a magnetic field with chaotic regions and islands and show the contours of the anisotropic diffusion equation reproduce key features in the field.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

STORM: Scrape-off layer turbulence in tokamak fusion reactors

The scrape-off layer of a tokamak fusion reactor carries the plasma exhaust from the hot core plasma to the material surfaces of the reactor vessel. The heat loads imposed by the exhaust are a critical limit on the performance of fusion power plants. Turbulent transport of the plasma regulates the width of the scrape-off layer plasma and must be modelled to understand the intensity of these heat loads. STORM is a plasma turbulence code capable of simulating three dimensional turbulence across the full scrape-off layer of a tokamak fusion reactor, using a drift reduced, collisional fluid model. STORM uses mostly finite difference schemes, with a staggered grid in the direction parallel to the magnetic field. We describe the model, geometry and initialisation options used by STORM, as well as the numerical methods, which are implemented using the BOUT++ plasma simulation framework. BOUT++ has been enhanced alongside the development of STORM, providing better support for staggered grid methods. We summarise these enhancements, including a detailed explanation of the parallel derivative methods, which underwent a major update for version 4 of BOUT++.

BOUT++

Bayesian Optimization for Anything (BOA): An open-source framework for accessible, user-friendly Bayesian optimization

We introduce Bayesian Optimization for Anything (BOA), a high-level Bayesian Optimization (BO) framework and model wrapping toolkit, which presents a novel approach to simplifying BO, with the goal of making it more accessible and user-friendly, particularly for those with limited expertise in the field. BOA addresses common barriers in implementing BO, focusing on ease of use, reducing the need for deep domain knowledge, and cutting down on extensive coding requirements. A notable feature of BOA is its language-agnostic architecture, which facilitates broader application in various fields and to a wider audience. We showcase BOA's application through three examples: a high-dimensional optimization with parameters of the SWAT+ watershed model, a highly parallelized optimization of this intrinsically non-parallel model, and a multi-objective optimization of the FETCH Tree-Crown Hydrodynamics model. Furthermore, these test cases illustrate BOA's effectiveness in addressing complex optimization challenges in diverse scenarios.

54 ENVIRONMENTAL SCIENCES

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING

Influence of rolling reduction and annealing on recrystallization and grain structure in Ta-2.5W alloys

The microstructural evolution of wrought Tantalum - 2.5 wt% Tungsten (Ta-2.5W) alloys during thermomechanical processing is critical for optimizing their mechanical reliability in demanding applications such as aerospace, chemical processing, and nuclear technology. Despite the widespread use of Ta-W alloys, a comprehensive understanding of how rolling reduction, annealing temperature, and elemental inhomogeneity interact to determine recrystallization behavior and grain refinement remains incomplete. Here, in this study, we systematically investigate the effects of cold rolling and subsequent annealing on the microstructure of Ta-2.5W, with particular attention to grain orientation, stored energy, and elemental banding. Our results demonstrate that higher rolling reduction rates lower the onset and completion temperatures for recrystallization, resulting in finer and more homogeneous grain structures. Electron backscatter diffraction (EBSD) analysis reveals that grains with a 〈111〉 parallel to the plate normal possess higher stored energy and nucleate recrystallization more readily than grains with a 〈001〉 parallel to the plate normal. Elemental mapping shows that tungsten inhomogeneity leads to localized bands of accelerated recrystallization and hardness variation. These findings provide new insights into the mechanisms of microstructure refinement in Ta-2.5W alloys, offering guidance for tailoring processing routes to achieve superior performance in demanding engineering environments.

Annealing