Search NASA⌕ Search

SEARCH · Search NASA

Results for “Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Exploiting Modern C++ for Portable Parallel Programming in Lattice QCD Applications

The evolution of ISO C++ standards increasingly serves the needs of scientific computing, offering potential benefits for developing portable applications. The recent revisions of C++ programming language, for instance, introduces a suite of algorithms capable of being executed on accelerators. Although this approach may not yield best performance, it can present a viable balance between code productivity and computational efficiency. In this report, we discuss the implementation of the HISQ operator utilizing a range of features from the C++17/20/23 standards and include an assessment of their performance.

Strelchenko, Alexei↗

A Portable Wave Tank and Wave Energy Converter for Engineering Dissemination and Outreach

Wave energy converters are a nascent energy generation technology that harnesses the power in ocean waves. To assist in communicating both fundamental and complex concepts of wave energy, a small-scale portable wave tank and wave energy converter have been developed. The system has been designed using commercial off-the-shelf components, and all design hardware and software are openly available for replication. This project builds on prior research conducted at Sandia National Laboratories, particularly in the areas of WEC device design and control systems. By showcasing the principles of causal feedback control and innovative device design, SIWEED not only serves as a practical demonstration tool but also enhances the educational experience for users. This paper presents the detailed system design of this tool. Furthermore, via testing and analysis, we demonstrate the basic functionality of the system.

educational↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

AthenaK: A Performance-portable Version of the Athena++ Adaptive Mesh Refinement Framework

We describe AthenaK: a new implementation of the Athena++ block-based adaptive mesh refinement framework using the Kokkos programming model. Finite volume methods for Newtonian, special relativistic, and general relativistic (GR) hydrodynamics and magnetohydrodynamics (MHD), and GR-radiation hydrodynamics and MHD, as well as a module for evolving Lagrangian tracer or charged test particles (e.g., cosmic rays) are implemented using the framework. In two companion papers, we describe (1) a new solver for the Einstein equations based on the Z4c formalism, and (2) a GRMHD solver in dynamical spacetimes also implemented using the framework, enabling new applications in numerical relativity. By adopting Kokkos, the code can be run on virtually any hardware, including CPUs, GPUs from multiple vendors, and emerging Advanced RISC Machine processors. AthenaK shows excellent performance and weak scaling, achieving over 1 billion cell updates per second for hydrodynamics in three dimensions on a single NVIDIA Grace Hopper processor. It does this with a typical parallel efficiency of 80% on 65,536 AMD GPUs on the OLCF Frontier system. Such performance portability enables AthenaK to leverage modern exascale computing systems for challenging applications in astrophysical fluid dynamics, numerical relativity, and multimessenger astrophysics.

79 ASTRONOMY AND ASTROPHYSICS↗

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as \ptor, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these mini-apps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

Atif, Mohammad [Brookhaven] (ORCID:000000026889770↗

Development of a Performance Portable Non-Equilibrium Plasma Fluid Solver on Adaptive Grids

This presentation will describe the numerical techniques, programming paradigms, verification, and performance of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures. Our plasma fluid model solves the conservation equations for self-consistent electrostatic Poisson, electron and heavy species transport, and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive mesh management library, AMReX (Zhang et al., JOSS, 4 (37) 1370, 2019), and can be built and run on widely available vendor specific GPU architectures (NVIDIA/AMD/Intel). We utilize a non-subcycled second order semi-implicit time-stepping method where all adaptive mesh refinement (AMR) levels are advanced with the same time step. The composite multi-level multigrid solver from within AMReX is used for each of the governing equations that are cast into a Helmholtz equation form. We have also developed a python based chemical mechanism parser framework that uses a similar format as CANTERA (Goodwin et al., Zenodo, 2018) yaml files as input. Our custom parser reads the yaml file and provides C++ files with transport and production rate functions that can be executed on both host (CPU) and device (GPU). We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on low-pressure capacitive and high-pressure streamer discharges. Our initial performance studies indicate 10X speed-up using 20 NVIDIA GPUs versus 200 CPUs for an atmospheric streamer discharge problem solved on a 512 x 1024 x 512 grid.

graphics processing units↗

Towards Portable Laser Desorption Postionization Mass Spectrometry Utilizing Microchip Laser Pulses

A 355 nm, 1.5 ns, up to 1 kHz microchip laser system has been tested to both yield ablation spots and postionize material within a plasma plume. A 30 µJ pulse energy microchip laser successfully ablated all samples: 6601 aluminum, 316 and 304 stainless steel, and Mo metal. While visible ablation craters were observed on all samples mentioned above, the stainless-steel samples displayed visible color changes in the resolidified iron, indicating ionization with the 355 nm pulses during ablation. Additionally, when used to intercept a plasma plume of Cs 2 CO 3 made with a separate laser system, the microchip laser pulses yielded Cs + emission signal otherwise not observed. With these experiments completed, the microchip laser has been successfully verified for readiness to be integrated into the front end of a mass spectrometer system as the ion source.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Portable Miniature Cryogenic Environment for In Situ Neutron Diffraction

Neutron diffraction instruments offer a platform for materials science and engineering studies at extended temperature ranges far from ambient. As one of the widely used neutron sample environment types, cryogenic furnaces are usually bulky and complex, and they may need hours of beamtime overhead for installation, configuration, cooling, and sample change, etc. To reduce the overhead time and expedite experiments at the state-of-the-art high-flux neutron source, we developed a low-cost, miniature, and easy-to-use cryogenic environment (77–473 K) for in situ neutron diffraction. A travel-size mug serves for the environment where the samples sit inside. Immediate cooling and an isothermal dwell at 77 K are realized on the sample by direct contact with liquid N 2 in the mug. The designed Al inserts serve as the holder of samples and heating elements, alleviate the thermal gradient, and clear neutron pathways. Both a single-sample continuous measurement and multi-sample high-throughput measurements are demonstrated in this environment. High-quality and refinable in situ neutron diffraction patterns are acquired on model materials. The results quantify the orthorhombic-to-cubic phase transformation process in LiMn 2 O 4 and differentiate the anisotropic lattice thermal expansions and bond length evolutions between rhombohedral perovskite oxides with composition variation.

47 OTHER INSTRUMENTATION↗

Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications

We characterize the GPU energy usage of two widely adopted exascale-ready applications representing two classes of particle and mesh solvers: (i) QMCPACK, a quantum Monte Carlo package, and (ii) AMReX-Castro, an adaptive mesh astrophysical code. We analyze power, temperature, utilization, and energy traces from double-/single (mixed)-precision benchmarks on NVIDIA’s A100 and H100 and AMD’s MI250X GPUs using queries in NVML and rocm_smi_lib, respectively. We explore application-specific metrics to provide insights on energy vs. performance trade-offs. Our results suggest that mixed-precision energy savings range between 6–25% on QMCPACK and 45% on AMReX-Castro. Also, we found gaps in the AMD tooling used on Frontier GPUs that need to be understood, while query resolutions on NVML have little variability between 1 ms-1 s. Overall, application level knowledge is crucial to define energy-cost/science-benefit opportunities for the codesign of future supercomputer architectures in the post-Moore era.

Godoy, William [ORNL] (ORCID:0000000225905178)↗

Deployment of portable, modular gas samplers as part of an atmospheric tracer experiment

Underground nuclear explosions release noble gases into the atmosphere that can be detected to support international monitoring efforts. Atmospheric transport models help predict the movement of these gases over long distances, but struggle to predict the movement in the atmosphere local to the release. A field experiment was designed to monitor the movement of 127 Xe within a 5-km radius. Four gas samplers were deployed as part of this experiment to collect atmospheric samples at various distances from the release point. In conclusion, these samples were then analyzed in a near-field lab using a NaI detector and in an off-site lab using gamma-gamma coincidence and beta-gamma coincidence counting.

Radioxenon↗

Evaluation of neutron dosimetry capabilities with the MC-15 portable multiplicity counter

This work proposes a preliminary neutron dose rate estimation method for a neutron multiplicity detector through measurement- and simulation-based analyses. Uncharacterized neutron-emitting sources may be encountered in situations such as nuclear emergency response, safeguards, and treaty verification. These circumstances may present irradiation risk to personnel conducting field assay, search, and characterization measurements. It is therefore of interest to provide a field neutron dosimetry capability with the existing neutron multiplicity counting (NMC) capabilities. To date, no commercially-available neutron detection systems are capable of both accurate NMC and real-time neutron dosimetry. This work will focus on estimating dose rate using input from a single fielded NMC called the MC-15. The energy-dependent neutron detection efficiency response of the MC-15 was quantified in monoenergetic neutron simulations and evaluated in response to two neutron-emitting sources and to a polyethylene-moderated source. The results were compared to existing neutron dosimeters and established the proof of concept for further investigation of the MC-15 for dose estimation. Measurement results were also replicated in simulations; additional simulations were then conducted to expand upon the limited empirical data. The initial empirical results provided a conversion factor appropriate for use when measuring 252 Cf neutrons that is independent of polyethylene shielding presence and thickness. The simulated data sets were then used to evaluate a fit equation allowing estimation of the neutron dose rate for less restricted geometric configurations and dependent only on the distance between the source and the detector. Additionally, the energy dependence of the efficiency response indicates that further empirical evaluations could provide energy-dependent conversion factors for broader neutron dosimetry capabilities with a wider range of neutron-emitting sources.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Portable Distortion Free Large Solid Angle Coverage Detector

Neutron single crystal diffractometers require large solid angle coverage for optimum performance. This can be achieved by tiling flat detectors in a cylindrical or spherical geometry, but results in large gaps in detector coverage and parallax distortion. The detector edges exhibit degraded resolution, distortion, and gamma rejection. Since the detector edge regions are a significant fraction of the detector active area, they must be removed from the experimental data set, requiring extra beam time to collect enough analysis data. A spherical detector with a continuous surface would effectively address this issue while eliminating most boundary ‘dead’ areas. Here we report on the development of a novel hemispherical shaped neutron detector using seamlessly tiled readout modules to form the desired shape. The heart of the detector is a specially developed curved neutron scintillator coupled to high resolution silicon photomultiplier (SiPM) Anger cameras via custom made fiber optic tapers (FOTs). The detector has been assembled and initial tests have been conducted at the High Flux Isotope Reactor (HFIR) beamlines at Oak Ridge National Laboratory (ORNL). Here, in this work, we describe details of the scintillator design, fabrication and characterization, evaluation of individual detector modules, the details of the detector design implementation, and evaluation of the assembled detector at ORNL beamlines.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

To Exascale and Beyond—The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM), a Performance Portable Global Atmosphere Model for Cloud-Resolving Scales

The new generation of heterogeneous CPU/GPU computer systems offer much greater computational performance but are not yet widely used for climate modeling. One reason for this is that traditional climate models were written before GPUs were available and would require an extensive overhaul to run on these new machines. In addition, even conventional “high–resolution” simulations don't currently provide enough parallel work to keep GPUs busy, so the benefits of such overhaul would be limited for the types of simulations climate scientists are accustomed to. The vision of the Simple Cloud-Resolving Energy Exascale Earth System (E3SM) Atmosphere Model (SCREAM) project is to create a global atmospheric model with the architecture to efficiently use GPUs and horizontal resolution sufficient to fully take advantage of GPU parallelism. After 5 years of model development, SCREAM is finally ready for use. In this paper, we describe the design of this new code, its performance on both CPU and heterogeneous machines, and its ability to simulate real-world climate via a set of four 40 day simulations covering all 4 seasons of the year.

54 ENVIRONMENTAL SCIENCES↗

Development of a portable high-T c spherical neutron polarimetry device at the Oak Ridge National Laboratory

Spherical neutron polarimetry is a powerful polarized neutron scattering technique used to determine complex magnetic structures which are only partly accessible by other methods. This technique measures the full neutron polarization change upon scattering from a sample by fully decoupling the incoming and outgoing neutron polarization with a zero-field chamber placed at the sample position. Recent advancements and testing are presented for a new spherical neutron polarimetry device utilizing high-T c superconducting YBCO films, PHiTPAD, at the High Flux Isotope Reactor at Oak Ridge National Laboratory. Furthermore, we introduce a conceptual design that utilizes wavelength-independent adiabatic transitions to adapt spherical neutron polarimetry for use with pulsed neutron sources, thereby expanding its potential applications in neutron scattering research.

47 OTHER INSTRUMENTATION↗

A new portable random number generator wrapper library

Random number generator is an important component of many scientific projects. Many projects are written using programming models (like OpenMP and SYCL) to target different architectures. However, some programming models do not provide a random number generator. In this work, we introduce our random number generator wrapper. It is a header-only library that supports three distributions of random numbers: uniform, normal, and poisson. On the GPU backend, it wraps the cuRAND and rocRAND library, and supports various random number engines. It also wraps random123, a counterbased random number generator, on both CPU and GPU. With this library, we can generate random numbers with a few lines of code and target both GPU and multi-thread CPU with the same code. We also investigate the performance and scalability of this wrapper on different architectures with different engines and the number of cores.

97 MATHEMATICS AND COMPUTING↗