Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Toward a 2D Local Implementation of Quantum Low-Density Parity-Check Codes

Geometric locality is an important theoretical and practical factor for quantum low-density parity-check (qLDPC) codes that affects code performance and ease of physical realization. For device architectures restricted to two-dimensional (2D) local gates, naively implementing the high-rate codes suitable for low-overhead fault-tolerant quantum computing incurs prohibitive overhead. In this work, we present an error-correction protocol built on a bilayer architecture that aims to reduce operational overheads when restricted to 2D local gates by measuring some generators less frequently than others. We investigate the family of bivariate-bicycle qLDPC codes and show that they are well suited for a parallel syndrome-measurement scheme using fast routing with local operations and classical communication (LOCC). Through circuit-level simulations, we find that in some parameter regimes, bivariate-bicycle codes implemented with this protocol have logical error rates comparable to the surface code while using fewer physical qubits. Published by the American Physical Society 2025

Berthusen, Noah (ORCID:0000000275862786)↗

Level-2 Milestone 9009: Flux and Rabbit Capabilities on El Capitan

This document is the milestone delivery report for the ASC 2025 L2 milestone (See Table 1) for advanced I/O capabilities for El Capitan via Flux Workload Manager support and the new I/O hardware designed for El Capitan, the Rabbit Storage System. In this document we describe the design of the Rabbit Storage System and how it is managed by Flux. We evaluate the performance and usability of Rabbit using ARES, IOR, and an AI inference workload. Overall, we find that Rabbit shows good scalability, especially in node-local storage configurations, and is more scalable than the global Lustre parallel file system.

97 MATHEMATICS AND COMPUTING↗

Enhancing ChatPORT with CUDA-to-SYCL Kernel Translation Capability

Large Language Models (LLMs) have shown strong capabilities in general code translation. However, code translation involving parallel programming models remains largely unexplored. This work enhances the capabilities of code LLMs in CUDA-to-SYCL kernel translation with parameter-efficient fine-tuning. The resultant fine-tuned LLM, called ChatPORT, is an effort to provide high-fidelity translations from one programming model to another. We describe the preparation of datasets from heterogeneous computing benchmarks for model fine-tuning and testing, the parameter-efficient fine-tuning of 19 open-source code models ranging in size from 0.5 to 34 billion parameters and evaluate the correctness rates of the SYCL kernels by the fine-tuned models. The experimental results show that most code models fail to translate CUDA codes to SYCL correctly. However, fine-tuning these models using a small set of CUDA and SYCL kernels can enhance the capabilities of these models in kernel translation. Depending on the sizes of the models, the correctness rate ranges from 19.9% to 81.7% for a test dataset of 62 CUDA kernels.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Fractional Skyrmion Tubes in Chiral‐Interfaced 3D Magnetic Nanowires

Magnetic skyrmions are chiral spin textures with rich physics and great potential for unconventional computing. Typically, skyrmions form in bulk crystals with reduced symmetry or ultrathin film multilayers involving heavy metals. Here, the formation of fractional Bloch skyrmion tubes at room temperature is demonstrated by 3D printing ferromagnetic double‐helix nanowires with two regions of opposite chirality. Using X‐ray microscopy and micromagnetic simulations, it is shown that the coexistence of vortex and anti‐parallel spin states induces the formation of fractional skyrmion tubes at zero magnetic fields, minimizing the energy cost of breaking the coupling between geometric and magnetic chirality. Control over zero‐field states is also demonstrated, including pure vortex, or mixed skyrmion‐vortex states, highlighting the magnetic reconfigurability of these 3D nanowires. This work shows how interfacing chiral geometries at the nanoscale can enable advanced forms of topological spintronics.

X-ray microscopy↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FunM2C: A Filter for Uncertainty Visualization of Multivariate Data on Multi-Core Devices

Uncertainty visualization is an emerging research topic in data visualization because neglecting uncertainty in visualization can lead to inaccurate assessments. In this paper, we study the propagation of multivariate data uncertainty in visualization. Although there have been a few advancements in probabilistic uncertainty visualization of multivariate data, three critical challenges remain to be addressed. First, the state-of-the-art probabilistic uncertainty visualization framework is limited to bivariate data (two variables). Second, existing uncertainty visualization algorithms use computationally intensive techniques and lack support for cross-platform portability. Third, as a consequence of the computational expense, integration into production visualization tools is impractical. In this work, we address all three issues and make a threefold contribution. First, we take a step to generalize the state-of-the-art probabilistic framework for bivariate data to multivariate data with an arbitrary number of variables. Second, through utilization of VTK-m’s shared-memory parallelism and cross-platform compatibility features, we demonstrate acceleration of multivariate uncertainty visualization on different many-core architectures, including OpenMP and AMD GPUs. Third, we demonstrate the integration of our algorithms with the ParaView software. We demonstrate the utility of our algorithms through experiments on multivariate simulation data with three and four variables.

Hari, Gautam↗

One-shot omnidirectional pressure integration through matrix inversion

In this work, we present a method to perform 2D and 3D omnidirectional pressure integration from velocity measurements with a single-iteration matrix inversion approach. This work builds upon our previous work, where the rotating parallel ray approach was extended to the limit of infinite rays by taking continuous projection integrals of the ray paths and recasting the problem as an iterative matrix inversion problem. This iterative matrix equation is now 'fast-forwarded' to the 'infinity' iteration, leading to a different matrix equation that can be solved in a single step, thereby presenting the same computational complexity as the Poisson equation. We observe computational speedups of ~10 6 when compared to brute-force omnidirectional integration methods, enabling the treatment of grids of ~10 9 points and potentially even larger in a desktop setup at the time of publication. Further examination of the boundary conditions of our one-shot method shows that omnidirectional pressure integration implements a boundary condition where the boundary points are treated as interior points to the extent that information is available. Finally, we show how the method can be extended from the regular grids typical of particle image velocimetry to the unstructured meshes characteristic of particle tracking velocimetry data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Linearised Fokker–Planck collision model for gyrokinetic simulations

We introduce a gyrokinetic, linearised Fokker–Planck collision model that satisfies conservation laws and is accurate at arbitrary collisionalities. The differential test-particle component of the operator is exact; the integral field-particle component is approximated using a spherical harmonic and a modified Laguerre polynomial expansion developed by Hirshman and Sigmar (1976 Phys. Fluids 19 1532). The numerical methods of the implementation in the δf-gyrokinetic code stella (Barnes et al 2019 J. Comput. Phys. 391 365–80) are discussed, and conservation properties of the operator are demonstrated. The collision model is then benchmarked against the collision model of the gyrokinetic solver GS2 in the limiting cases of a reduced test-particle collision operator and energy- and momentum-conserving operator. The accuracy of the full collision model is investigated by solving the parallel Spitzer-Härm problem for the transport coefficients. It is shown that retaining collisional energy flux and higher-order terms in the field-particle operator reduces errors in the transport coefficients from 10%–25% for a simple momentum- and energy-conserving model to under 1%.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enriched immersed finite element and isogeometric analysis: algorithms and data structures

Immersed finite element methods provide a convenient analysis framework for problems involving geometrically complex domains, such as those found in topology optimization and microstructures for engineered materials. However, their implementation remains a major challenge due to, among other things, the need to apply nontrivial stabilization schemes and generate custom quadrature rules. This article introduces the robust and computationally efficient algorithms and data structures comprising an immersed finite element preprocessing framework. The input to the preprocessor consists of a background mesh and one or more geometries defined on its domain. The output is structured into groups of elements with custom quadrature rules formatted such that common finite element assembly routines may be used without or with only minimal modifications. The key to the preprocessing framework is the construction of material topology information, concurrently with the generation of a quadrature rule, which is then used to perform enrichment and generate stabilization rules. While the algorithmic framework applies to a wide range of immersed finite element methods using different types of meshes, integration, and stabilization schemes, the preprocessor is presented within the context of the extended isogeometric analysis. This method utilizes a structured B-spline mesh, a generalized Heaviside enrichment strategy considering the material layout within individual basis functions’ supports, and face-oriented ghost stabilization. Using a set of examples, the effectiveness of the enrichment and stabilization strategies is demonstrated alongside the preprocessor’s robustness in geometric edge cases. Additionally, the performance and parallel scalability of the implementation are evaluated.

Computer implementation↗

Computing material volume fractions on a superimposed mesh as applied to Monte Carlo particle transport simulations

Here, we present a newly implemented ray tracing algorithm in OpenMC for efficiently computing material volume fractions on superimposed meshes in complex geometries. By firing rays along each coordinate direction through the geometry, the approach accumulates track-length data in each mesh element, thereby determining the fractional composition of each material. Scaling studies on three different models—a random tetrahedra configuration, the Frascati Neutron Generator ITER dose rate benchmark, and a stellarator design—show excellent parallel performance, with nearly linear speedup on modern multi-threaded and distributed-memory systems. An analysis of the residual error relative to high-resolution reference solutions demonstrated that under optimal conditions it decreases as 1/R, where R is the number of rays fired, making it straightforward to achieve user-prescribed accuracy. This new functionality enables practical, mesh-based approaches for detailed nuclear analyses in production Monte Carlo workflows without resorting to expensive, fully conformal or unstructured meshing.

Monte Carlo↗

Multi-GPU porting of a phase-change cascaded lattice Boltzmann method for three-dimensional pool boiling simulations

The Lattice Boltzmann method (LBM) has proven effective in simulating phase-change phenomena, such as melting, solidification, evaporation, and boiling. In this work, we develop a highly parallelized multi-GPU implementation of LBM for three-dimensional pool boiling simulations. The code is based on the OpenACC programming model, which enables the code to be deployed efficiently on multi-core CPUs, GPUs, and potentially other accelerators, without the need for architecture-specific rewrites. To support large-scale simulations, the domain is decomposed and distributed across multiple compute nodes using MPI. We demonstrate that the code exhibits excellent scaling properties, with ideal strong-scaling running with up to 256 GPUs on the MareNostrum5 cluster.

97 MATHEMATICS AND COMPUTING↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

Recorder trace files of 91 built-in tests from three widely-used I/O libraries (Dec 24, 2024)

This dataset contains the Recorder trace files of 91 built-in tests from three widely-used I/O libraries. It was utilized in our IPDPS'25 paper, "VerifyIO: Verifying Adherence to Parallel I/O Consistency Semantics". The dataset can be used to reproduce the results presented in the paper and to conduct additional analyses

Wang, Chen [Lawrence Livermore National Laboratory↗

Towards modelling AR Sco: calibration – reproducing high-energy pulsar emission and testing convergence to Aristotelian electrodynamics

In recent years, kinetic simulations have been crucial to further our understanding of pulsar electrodynamics. Yet, due to the large-scale separation between the gyro-period and the stellar rotation period, resolving the particle gyration has been computationally unfeasible for realistic pulsar parameters. The main aim of this work is comparing our gyro-phase-resolved model with a gyro-centric pulsar model, where our model solves the general equations of motion with included radiation reaction using a higher order numerical solver with adaptive time-steps. Specifically, we aim to (i) reproduce a pulsar’s high-energy emission maps, namely one with 10 per cent of the surface B-field strength of Vela, and the spectra produced by an independent gyro-centric pulsar emission model; and (ii) test convergence of these results to the radiation-reaction limit of Aristotelian electrodynamics. (iii) Additionally, we identify the effect that a large $E_{\parallel }$-field has on the trajectories and radiation calculations. We find that we can reproduce the curvature radiation emission maps and spectra well, using 10 per cent field strengths of the Vela pulsar and injecting our particles at a higher altitude in the magnetosphere. Using sufficiently large $E_{\parallel }$-fields, our numeric results converge to the analytic radiation-reaction limit trajectories. Additionally, we illustrate the importance of accounting for the $\mathbf {E}\times \mathbf {B}$-drift in the particle trajectories and radiation calculations, validating the Harding and collaborators’ model approach. Lastly, we found that our model deals very well with the high-radiation-reaction and high-field regimes present in pulsars.

79 ASTRONOMY AND ASTROPHYSICS↗

Electron density measurements and calculations in a helium capacitively-coupled radio-frequency plasma

We report a comparison of inferred electron density (n e ) in a He capacitively-coupled plasma, deduced from laser-collision induced fluorescence measurements, with values computed using a hybrid simulation framework based on particle-in-cell/Monte Carlo collisions simulations and a fluid model for excited He atoms. The studies were carried out for gas pressures between 50 mTorr and 1000 mTorr and peak-to-peak radio-frequency (13.56 MHz) voltages between 150 V and 350 V, in a highly symmetric source equipped with plane-parallel electrodes. A good agreement is found between the experimental and modeling results for n e except at the lowest operating voltages and gas pressures. The (effective) electron temperature (T e ) values derived by the two methods agree as well reasonably within the plasma bulk. The simulation results are used to compare the density distributions of He + and various He excited levels and their major populating and de-populating channels at 100 mTorr and 1000 mTorr.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Computational Analysis of Hydraulic Efficiency of Michigan DOT Cover C

Drainage structures are used to capture stormwater runoff in streets and highways in urban environments. These drainage structures, which typically consist of catch basins with grates, inlets, or combination grates/inlets, collect stormwater runoff and discharge through buried conveyance systems. They are strategically placed for public safety in curb and gutter systems to provide efficient drainage of water from roadways and thus reduce the risk of hydroplaning. The performance of drainage structures is measured in terms of hydraulic efficiency, which is defined as the percentage of flow captured by the basin as compared to the total flow drainage to the structure. Understanding of the performance of these drainage structures helps designers properly space inlets to promote an economic design that ensures the safety of the traveling public. The current design methodology used to determine drainage structure follows guidelines established in the current Michigan Department of Transportation (MDOT) Drainage Manual (2006). The guidelines in the MDOT Drainage Manual were modeled after the Federal Highway Administration’s (FHWA) Hydraulic Engineering Circular 22 “Urban Drainage Design” (HEC-22). HEC-22 includes empirically derived equations to calculate the interception capacity of drainage structures for several commonly used grate configurations, such as the parallel bar, curved vane and tilt bar grates, which are based on a research study performed by Burgi et al. in the 1970s. MDOT uses several drainage structures to capture runoff that are detailed as Standard Plans. Many of these drainage structures utilize sinusoidal type grates that are not described in HEC-22. Physical modeling of these structures has been limited, posing the need to have them analyzed to verify their capture efficiency. Current MDOT practice is to assume a similar sized reticuline grate, as described in HEC-22, for capture efficiencies. Until recently, evaluating the hydraulic performance of drainage structures was limited to physical modeling in a hydraulics laboratory. With advances in engineering software and computing power, computational fluid dynamics (CFD) modeling has become a more cost-effective alternative. The Federal Highway Administration (FHWA) provides states the option to evaluate their drainage structures using CFD through the Transportation Pooled Fund Program. This study, “Computational Analysis of Hydraulic Efficiency of Michigan DOT Cover C,” was carried out using the pooled fund. MDOT’s Cover C was chosen as the first test candidate, given its similar sinusoidal pattern to other MDOT grates, but it is typically used for high-volume, higher speed applications. A similar version, Cover CX, is used on interstate highways but does not have traverse bars for bicycle safety. Additional grates may be considered for evaluation in the future.

42 ENGINEERING↗

Tough Errors are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This project builds toward a comprehensive error-mitigating toolkit that makes quantum programming more robust and adaptive to the noisy, resource-limited nature of today’s quantum hardware. To that end, it integrates established error-mitigation methods — such as zero-noise extrapolation and dynamical decoupling — directly into compiler infrastructures. These techniques will be packaged as modules that can automatically adjust and combine based on performance analysis, enabling compilers to explore large design spaces and produce optimized, low-noise quantum programs with minimal manual intervention. In parallel, this project also explores new approaches to analog quantum programming or quantum simulation, and has developed the programming language SimuQ which treats quantum Hamiltonian evolution as the central object.

97 MATHEMATICS AND COMPUTING↗