Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

Reconfigurable unitary transformations of optical beam arrays

Spatial transformations of light are ubiquitous in optics, with examples ranging from simple imaging with a lens to quantum and classical information processing in waveguide meshes. Multi-plane light converter (MPLC) systems have emerged as a platform that promises completely general spatial transformations, i.e., a universal unitary. However, until now, MPLC systems have demonstrated transformations that are far from general, e.g., converting from a Gaussian to Laguerre-Gauss mode. Here, we demonstrate the promise of an MLPC, the ability to impose an arbitrary unitary transformation that can be reconfigured dynamically. Specifically, we consider transformations on superpositions of parallel free-space beams arranged in an array, which is a common information encoding in photonics. We experimentally test the full gamut of unitary transformations for a system of two parallel beams and make a map of their fidelity. We obtain an average transformation fidelity of 0.85 ± 0.03. This high-fidelity suggests that MPLCs are a useful tool for implementing the unitary transformations that comprise quantum and classical information processing.

47 OTHER INSTRUMENTATION↗

Mode-multiplexed photonic integrated vector dot-product core from inverse design

Photonic computing has the potential to harness the full degrees of freedom (DOFs) of the light field, including the wavelength, spatial mode, spatial location, phase quadrature, and polarization, to achieve a higher level of computing parallelism and scalability than digital electronic processors. While multiplexing using the wavelength and other DOFs can be readily integrated on silicon photonics platforms with compact footprints, conventional mode-division multiplexed (MDM) photonic designs occupy areas exceeding tens to hundreds of microns for a few spatial modes, significantly limiting their scalability. Here, we utilize inverse design to demonstrate an ultracompact photonic computing core that calculates vector dot products based on MDM coherent mixing. Our dot-product core integrates the functionalities of two-mode multiplexers and one multimode coherent mixer within a nominal footprint of 5 μm x 3 μm . We have experimentally demonstrated computing examples on the fabricated dot-product core, including complex number multiplication and motion estimation using optical flow. The compact dot-product core design enables large-scale on-chip integration in a parallel photonic computing primitive cluster for high-throughput scientific computing and computer vision tasks.

97 MATHEMATICS AND COMPUTING↗

High-Fidelity Modeling of a Type-5 Wind Turbine Gearbox (Intern Poster) [Poster]

Type-5 wind turbines are unique in their use of a permanent magnet synchronous generator, as well as their use of a hydraulic torque converter. This architecture presents an opportunity to provide steady and grid-ready energy without the need for a power converter. With infrastructure continuity and reliability being an important topic amongst renewable energies, researchers have been prompted to further investigate the benefits of type-5 turbines’ unique electromechanical configuration on stable electricity generation. Researchers involved in the WindSG project, SG standing for synchronous generator, are aiming to model a type-5 turbine using Real Time Digital Simulation (RTDS) to evaluate its efficacy in the grid. RSCAD, the software run on the RTDS, comes pre-loaded with electrical and electromechanical components to help simulate electrical generation and grid conditions. However, within this repertoire there is a lack of a component to represent a gearbox with high-fidelity. Within RSCAD’s case studies, the gearbox is often represented simply by a gear ratio value. This presented the task of developing a high-fidelity gearbox model in RSCAD for use in the larger RTDS type-5 wind turbine model. This poster describes a method of developing a lumped parameter mathematical model to represent a planetary-parallel-parallel gearbox in RSCAD for use in RTDS.

17 WIND ENERGY↗

The VTK-m User's Guide (V. 2.2)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗

High-Fidelity Modeling of a Type-5 Wind Turbine Gearbox (Intern Technical Presentation) (Poster)

Type-5 wind turbines are unique in their use of a permanent magnet synchronous generator, as well as their use of a hydraulic torque converter. This architecture presents an opportunity to provide steady and grid-ready energy without the need for a power converter. With infrastructure continuity and reliability being an important topic amongst renewable energies, researchers have been prompted to further investigate the benefits of type-5 turbines’ unique electromechanical configuration on stable electricity generation. Researchers involved in the WindSG project, SG standing for synchronous generator, are aiming to model a type-5 turbine using Real Time Digital Simulation (RTDS) to evaluate its efficacy in the grid. RSCAD, the software run on the RTDS, comes pre-loaded with electrical and electromechanical components to help simulate electrical generation and grid conditions. However, within this repertoire there is a lack of a component to represent a gearbox with high-fidelity. Within RSCAD’s case studies, the gearbox is often represented simply by a gear ratio value. This presented the task of developing a high-fidelity gearbox model in RSCAD for use in the larger RTDS type-5 wind turbine model. This presentation describes a method of developing a lumped parameter mathematical model to represent a planetary-parallel-parallel gearbox in RSCAD for use in RTDS.

17 WIND ENERGY↗

Assessment of Current MACCS Capabilities for Modeling Atmospheric Physical and Chemical Transformations

The physical and chemical transformation during atmospheric transport of radionuclides released into the environment has the possibility of impacting consequence modeling results. Accordingly, this report identifies physical and chemical transformations that may occur following release of chemically reactive radioactive species, how those transformations may affect modeling of consequences of release to the atmosphere and identifies current capabilities – in both MACCS and other state-of-practice atmospheric transport and dispersions models– to model those transformations. It was found that the inclusion of physical and chemical transformations is currently very limited in current state-of-practice codes for atmospheric dispersion of radionuclides. State-of-practice atmospheric dispersion codes appear to be typically limited to simulating either physical-chemical transformations or radioactive transformations, but not both. A state-of-practice atmospheric dispersion code capable of performing parallel physical, chemical, and radioactive transformation was not identified. A few atmospheric dispersion codes capable of modeling physical and chemical transport of specific species such as tritium or uranium hexafluoride were identified. Consequently, there is currently no information available that clearly suggests updates to the MACCS code are needed to bring it up to state-of-practice. However, investigations concluded that the MACCS computational framework can currently accommodate multiple physical/chemical forms in one simulation. Additionally, with some major assumptions, the computational framework in MACCS can accommodate parallel physical-chemical and radioactive transformations.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Portable HCAL reconstruction in the CMS detector using the Alpaka library

CMS has deployed a number of different GPU algorithms at the High-Level Trigger (HLT) in Run 3. As the code base for GPU algorithms continues to grow, the burden for developing and maintaining separate implementations for GPU and CPU becomes increasingly challenging. To mitigate this, CMS has adopted the Alpaka (Abstraction Library for Parallel Kernel Acceleration) library as the performance portability solution to provide a single-code base for parallel execution on both GPUs and CPUs in CMS software (CMSSW). A direct CUDA version of HCAL energy reconstruction, called Minimization At Hcal, Iteratively (MAHI), has been deployed at the HLT in the 2022-2023 data taking period. This contribution will describe how the CUDA version is converted into a portable implementation using the Alpaka library. We will discuss the porting experience from CUDA to Alpaka, the validation process and the performance of the Alpaka version in CPU and GPU.

Kwok, Martin↗

Sierra/SD – User's Manual – 5.22

Sierra/SD provides a massively parallel implementation of structural dynamics finite element analysis, required for high-fidelity, validated models used in modal, vibration, static and shock analysis of weapons systems. This document provides a user’s guide to the input for Sierra/SD. Details of input specifications for the different solution types, output options, element types and parameters are included. The appendices contain detailed examples, and instructions for running the software on parallel platforms.

97 MATHEMATICS AND COMPUTING↗

Thermonuclear Burn in a Multiphysics Code on GPUs

Multiphysics codes links to a library called SINGE for the calculation of thermonuclear (TN) burn rates, but some current multiphysics codes do not attempt to leverage the support for parallel operation that SINGE provides. Our goal is to investigate implementations of the SINGE workflow and analyze how the use of a performance portability layer could reduce run time on CPU archi tectures while also supporting GPU architectures without requiring code modifications. We looked to the Kokkos C++ Performance Portability Ecosystem to implement hardware agnostic parallel patterns.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ensemble Simulation Techniques and Fast Randomized Algorithms

The major goals of the project were to develop and analyze new ensemble simulation techniques, including trajectory stratification and preconditioned MCMC techniques, as well as develop fast numerical linear algebra techniques closely related to ensemble simulation ideas. The trajectory stratification techniques involve simulating in parallel short trajectory fragments of a Markov process confined to a specific region of space‐time and then patching together the statistics gathered to assemble estimates of very general dynamical properties. We have also developed this approach for rare event simulation and extended the techniques to applications requiring a more general framework (such as electronic structure calculations). The preconditioned MCMC techniques involve simulating multiple Markov chains in parallel and then using information from the ensemble to speed the mixing of each individual chain. The fast randomized linear algebra methods are motivated by the diffusion Monte Carlo technique, but are applicable to finding the dominant eigenvalue of (almost) general matrices. For most non‐negative matrices, the schemes result in an error (compared to the power method) that is constant in the dimension of the problem. For more general matrices, we see a very clear sublinear cost trend in computational tests.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Sierra/SD – User's Manual (V.5.24)

Sierra/SD provides a massively parallel implementation of structural dynamics finite element analysis, required for high-fidelity, validated models used in modal, vibration, static and shock analysis of weapons systems. This document provides a user’s guide to the input for Sierra/SD. Details of input specifications for the different solution types, output options, element types and parameters are included. The appendices contain detailed examples, and instructions for running the software on parallel platforms.

42 ENGINEERING↗

Process–Property–Performance Mapping of Additively Manufactured 316H Stainless Steel Components

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development of advanced materials and components fabricated via additive manufacturing, and is using laser powder bed fusion (LPBF) of 316H stainless steel as an initial case study. In the previous fiscal year, miniature high-throughput specimens were printed on multiple LPBF systems to provide initial processing windows to minimize porosity and limit epitaxial grain growth during prints. This fiscal year, scaled builds were completed on three different LPBF systems at ORNL: a GE Concept Laser M2, a Renishaw AM400, and an EOS M290. Builds on the Concept Laser were conducted on multiple powder lots and processing parameter ranges to provide microstructure effects on time-independent and time-dependent mechanical properties. Builds on the Renishaw were produced using Oak Ridge National Laboratory (ORNL)-optimized printing parameters and Argonne National Laboratory (ANL)-optimized printing parameters to compare outcomes of parallel process optimization efforts at different national laboratories on the same LPBF system. Similarly, the build completed on the EOS M290 replicated the processing parameters of builds completed at Los Alamos National Laboratory (LANL). Optical microscopy and electron backscatter diffraction characterization was completed on all builds. In addition to the general round robin characterization, this work-package generated time-independent data, including tensile and fracture toughness test data on scaled Concept Laser builds as a function of processing parameters and post-build heat treatment. This analysis is complimentary to work in parallel work packages aiming to establish heat treatment and processing effects on time-dependent properties. It was found that although the stress-relief heat treatment provides the highest strength at lower-temperatures, tensile strength begins to converge at higher temperatures regardless of heat treatment condition. In addition, the more rigorous solution annealing and hot-isostatic pressing post-build heat treatments result in higher fracture toughness than the stress-relieved condition. The root-causes of the lower fracture toughness of the stress-relieved LPBF 316H material was informed via a stress-relief optimization study on a scaled concept laser print, where it was found that although dislocation recovery was largely complete after only a couple hours at 650°C, the extended hold of the current 24h heat treatment employed on scaled builds likely caused increased carbide volume fractions along the LPBF 316H grain boundaries, thereby deteriorating crack propagation resistance. This trend was seen to become more deleterious with additional increases of stress-relief temperature to 750°C or 850°C. These results have helped inform a new optimal stress-relief annealing condition for LPBF 316H for future campaign testing (650°C for 2h).

36 MATERIALS SCIENCE↗

Development of Transformative Preparation Methods to Push up High Q&G Performance of FRIB Spare HWR Cryomodule Cavities

The FRIB accelerator project construction, a top priority of US nuclear science, was completed in January 2022, and is now moving to user operation. The stable and reliable operation of the accelerating cryomodules is essential in achieving/fulfilling DOE and user expectations. So far, FRIB cryomodules meet all FRIB specifications for cavity performance. However, during the lifetime of machine operation, degradation of cryomodule performance is possible, as reported in similar operating facilities (CEBAF, SNS). If cryomodule degradation is observed at FRIB, the under-performing cryomodule will require replacement/maintenance. In effort to manage operational reliability, FRIB plans to construct a 0.53 half-wave cryomodule to serve as an active spare. In a parallel effort, FRIB will also work toward increasing operational Q and gradient of spare cryomodule cavities to gain an overall performance margin to support future operational reliability. The current FRIB cavity designs have a potential to operate at gradients higher than 8 MV/m, but are currently limited by field emission (FE) and/or high field Q slope (HFQS); known issue in buffered chemical polished (BCP) treated cavities. The proposal looks to develop transformative surface preparation treatments to improve the operational gradient of spare cryomodules higher than 10 MV/m while maintaining high Q. Thus, increasing operational margin by 30 - 50%. With the goal to improve operational reliability set, the proposal will investigate multiple objectives as possible paths forward to achieve an overall increase in cavity performance and gain a better understanding of SRF limiting mechanisms. The proposal will study the application of different chemical surface treatments to 0.53 half-wave cavities, with the addition of low temperature bakes (LTB), and measure their effects on accelerating performance. Proposed chemical treatments to be explored in this proposal include conventional EP acid mixtures, as well as innovated EP and BCP acid mixtures designed to simplify processing paths in migrating FE and HFQS. The proposed transformative treatment wet N-doping also has the potential to replicate recent advancements in SRF technology relating to nitrogen doping and high Q operation without the requirement for an ultra-high vacuum annealing furnace; currently being developed at FNAL and JLAB. In parallel, high Q performance relating to flux trapping will be investigated with the installation of a second layer of magnetic shielding in the vertical test Dewar. The research objectives presented in the proposal, and their corresponding effects on cavity performance, will provide essential knowledge and future guidance to the SRF community and provide possible paths for future SRF based projects and applications.

43 PARTICLE ACCELERATORS↗

Conceptual Designs for Irradiation Creep Testing of SiC in HFIR

Understanding irradiation creep of nuclear fuel cladding is important to properly size the initial fuel-cladding gap and understand when pellet-cladding contact is expected to occur due to a combination of fuel swelling and cladding creep-down. Irradiation creep also plays a role in relaxing stresses that develop in-pile. Silicon carbide fiber–reinforced silicon carbide matrix (SiC/SiC) composites are the leading long-term accident-tolerant fuel cladding concept for light-water reactors (LWRs). Although some limited data are available regarding irradiation creep of the individual constituents (fibers, matrix), data regarding irradiation creep of SiC/SiC composites are currently insufficient. Additional data regarding irradiation creep compliance and the rupture lifetime (combination of creep and slow crack growth) are needed to understand material limitations. This work describes the design and development of two irradiation vehicles that are being pursued for testing SiC/SiC concepts in the High Flux Isotope Reactor (HFIR). The first is a passive experiment, referred to as the PRECISE experiment, that leverages the constant coolant pressure of HFIR to compress a metallic bellows and provide a well-characterized load to drive creep in a SiC/SiC dog bone specimen. The total creep strain would be quantified post-irradiation by measuring dimensional changes of the specimen length as well as local dimensional changes within the gauge region. Non-stressed specimens would also be irradiated under the same conditions to provide an indication of dimensional changes due to radiation-induced swelling in the absence of creep. A second, more complex experiment, referred to as the INSITE experiment, is being designed in parallel that would use pneumatics to pressurize a metal bellows and linear variable differential transformers (LVDTs) to measure the specimen displacement in situ during irradiation. Such an experiment would provide significantly more data regarding the evolution of the creep compliance as a function of dose and applied stress within a single experiment but would require significantly more development time and cost to execute. The primary concern with the INSITE experiment is the accuracy, reliability, and expected lifetime of the LVDTs during irradiation at elevated temperatures. Efforts are being made to adjust the experiment design and operating procedure to limit LVDT temperatures and mitigate or otherwise compensate for uncertainties due to factors such as temperature fluctuations, creep in the surrounding structural materials, and drift of the LVDTs. This work describes the experiment designs, thermal and structural analysis that were performed to ensure that the desired temperature and stress conditions can be achieved, some initial sensitivity analyses to predict the evolution of the radiation-induced specimen displacements, and potential sources of uncertainty in the measurements. Out-of-pile testing is being performed in parallel to confirm that the test trains achieve the expected stress states in the specimens and do not result in prohibitive stress concentrators (e.g., in the grip regions) that might risk pre-mature failure. The PRECISE experiments are proceeding toward fabrication and assembly with HFIR insertion planned during fiscal year 2026. The INSITE experiment is progressing toward out-of-pile demonstrations, which will provide more conclusive evidence regarding the feasibility of executing these tests in HFIR or whether alternative displacement monitoring techniques may need to be considered.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

The Viskores User's Guide (V.1.0)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created Viskores: the visualization toolkit for multi-/many-core architectures. Viskores supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. Viskores also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although Viskores provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗