Search NASA⌕ Search

SEARCH · Search NASA

Results for “Distributed computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Charged particle transport coefficient challenges in high energy density plasmas

High energy density physics (HEDP) and inertial confinement fusion (ICF) research typically relies on computational modeling using radiation-hydrodynamics codes in order to design experiments and understand their results. These tools, in turn, rely on numerous charged particle transport and relaxation coefficients to account for laser energy absorption, viscous dissipation, mass transport, thermal conduction, electrical conduction, non-local ion (including charged fusion product) transport, non-local electron transport, magnetohydrodynamics, multi-ion-species thermalization, and electron-ion equilibration. In many situations, these coefficients couple to other physics, such as imposed or self-generated magnetic fields. Furthermore, how these coefficients combine are sensitive to plasma conditions as well as how materials are distributed within a computational cell. Uncertainties in these coefficients and how they couple to other physics could explain many of the discrepancies between simulation predictions and experimental results that persist in even the most detailed calculations. This paper reviews the challenges faced by radiation-hydrodynamics in predicting the results of HEDP and ICF experiments with regard to these and other physics models typically included in simulation codes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Desmearing two-dimensional small-angle neutron scattering data by central moment expansions

Resolution smearing is a critical challenge in the quantitative analysis of two-dimensional small-angle neutron scattering (SANS) data, particularly in studies of soft-matter flow and deformation using SANS. Here, we present a central moment expansion technique to address smearing in anisotropic scattering spectra, offering a model-free desmearing methodology. By accounting for directional variations in resolution smearing and enhancing computational efficiency, this approach reconstructs desmeared intensity distributions from smeared experimental data. Computational benchmarks using interacting hard-sphere fluids and Gaussian chain models validate the accuracy of the method, while simulated noise analyses confirm its robustness under experimental conditions. Experimental validation using rheological SANS data from shear-induced micellar structures demonstrates the practicality and effectiveness of the proposed algorithm. The desmearing technique provides a powerful tool for advancing the quantitative analysis of anisotropic scattering patterns, enabling precise insights into the interplay between material microstructure and macroscopic flow behavior.

anisotropic scattering spectra↗

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Object Proxy Patterns for Accelerating Distributed Applications

Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.

Distributed Computing↗

Computing an Optimal Entanglement Path with Throughput and Fidelity Considerations

Entanglement distribution is a core function of quantum networks essential for operations including teleportation, distributed quantum sensing, and multisite computation. Entanglement throughput and fidelity are two critical performance measures that depend on the quantum transmission along the links and swapping operations at the repeaters along the path. We study the problem of computing a end-to-end entanglement path that satisfies both fidelity and throughput requirements, leveraging qubit buffers at the nodes and considering the sequential swapping order. We show that the general problem of simultaneously satisfying both metrics to be NP-hard, and develop an algorithm to maximize throughput subject to a given fidelity threshold. We introduce the concepts of entanglement probability distribution and path domination and exploit them in the design of our algorithm. Extensive numerical results show that our algorithm can find optimal solutions in networks with thousands of nodes in less than a second. We also describe practical and possible implementation aspects of this algorithm in terms of devices and architecture support.

Xue, Guoliang [Arizona State University]↗

Multi-GPU porting of a phase-change cascaded lattice Boltzmann method for three-dimensional pool boiling simulations

The Lattice Boltzmann method (LBM) has proven effective in simulating phase-change phenomena, such as melting, solidification, evaporation, and boiling. In this work, we develop a highly parallelized multi-GPU implementation of LBM for three-dimensional pool boiling simulations. The code is based on the OpenACC programming model, which enables the code to be deployed efficiently on multi-core CPUs, GPUs, and potentially other accelerators, without the need for architecture-specific rewrites. To support large-scale simulations, the domain is decomposed and distributed across multiple compute nodes using MPI. We demonstrate that the code exhibits excellent scaling properties, with ideal strong-scaling running with up to 256 GPUs on the MareNostrum5 cluster.

97 MATHEMATICS AND COMPUTING↗

High-Fidelity Multiphysics Modeling of a Heat Pipe Microreactor Using BlueCrab

Researchers who are actively developing nuclear microreactors are planning to employ innovative designs and features using traditional commercial modeling tools that may be inadequate for their design and licensing activities. The codes developed under the U.S. Department of Energy Office of Nuclear Energy Advanced Modeling and Simulation (NEAMS) program provide flexibility in terms of geometry modeling and multiphysics coupling and are particularly well suited for modeling novel microreactor concepts. To test the maturity of these codes, this paper introduces a conceptual heat pipe microreactor (HP-MR) designed to gather various technologies of interest to microreactor developers such as control drums, heat pipes, and hydride moderators. Here, the objective of this effort is to demonstrate NEAMS tools capability to perform high-fidelity multiphysics simulations, using coupled neutronics (via the Griffin code), heat conduction (via the BISON code), heat pipe modeling (via the Sockeye code), and hydrogen redistribution in hydride metal moderator (via the SWIFT code). Codes are coupled in-memory through the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, which permits flexible multiphysics data transfer schemes. The analysis confirmed two key aspects of the HP-MR concept: (1) its ability to follow the power load requested from the heat pipe and (2) its ability to avoid heat pipe cascading failure unless designed with high power close to operating failure limits of its heat pipes. The developed computational model was distributed publicly on the Virtual Test Bed for training purposes to accelerate adoption by industry and to provide a high-fidelity multiphysics solution for benchmarking against other tools. Additional multiphysics analyses including other transients and coupled physics were identified as necessary future work, together with a focus on validating multiphysics behavior against experiments.

Microreactor↗

Physics-inspired spatiotemporal-graph AI ensemble for the detection of higher order wave mode signals of spinning binary black hole mergers

We present a new class of AI models for the detection of quasi-circular, spinning, non-precessing binary black hole mergers whose waveforms include the higher order gravitational wave modes ($\ell$, |m|) = {(2,2), (2,1), (3,3), (3,2), (4,4)}, and mode mixing effects in the $\ell$ = 3, |m| = 2 harmonics. These AI models combine hybrid dilated convolution neural networks to accurately model both short- and long-range temporal sequential information of gravitational waves; and graph neural networks to capture spatial correlations among gravitational wave observatories to consistently describe and identify the presence of a signal in a three detector network encompassing the Advanced LIGO and Virgo detectors. We first trained these spatiotemporal-graph AI models using synthetic noise, using 1.2 million modeled waveforms to densely sample this signal manifold, within 1.7 h using 256 NVIDIA A100 GPUs in the Polaris supercomputer at the Argonne Leadership Computing Facility. This distributed training approach exhibited optimal classification performance, and strong scaling up to 512 NVIDIA A100 GPUs. With these AI ensembles we processed data from a three detector network, and found that an ensemble of 4 AI models achieves state-of-the-art performance for signal detection, and reports two misclassifications for every decade of searched data. We distributed AI inference over 128 GPUs in the Polaris supercomputer and 128 nodes in the Theta supercomputer, and completed the processing of a decade of gravitational wave data from a three detector network within 3.5 h. Finally, we fine-tuned these AI ensembles to process the entire month of February 2020, which is part of the O3b LIGO/Virgo observation run, and found 6 gravitational waves, concurrently identified in Advanced LIGO and Advanced Virgo data, and zero false positives. This analysis was completed in one hour using one NVIDIA A100 GPU.

79 ASTRONOMY AND ASTROPHYSICS↗

Boosting H I -Galaxy Cross-Clustering Signal through Higher-Order Cross-Correlations

After reionization, neutral hydrogen (${\rm H\, \small {I}}$) traces the large-scale structure (LSS) of the Universe, enabling ${\rm H\, \small {I}}$ intensity mapping (IM) to capture the LSS in 3D and constrain key cosmological parameters. We present a new framework utilizing higher-order cross-correlations to study ${\rm H\, \small {I}}$ clustering around galaxies, tested using real-space data from the IllustrisTNG300 simulation. This approach computes the joint distributions of k-nearest neighbor (kNN) optical galaxies and the ${\rm H\, \small {I}}$ brightness temperature field smoothed at relevant scales (the kNN-field framework), providing sensitivity to all higher-order cross-correlations, unlike two-point statistics. To simulate ${\rm H\, \small {I}}$ data from actual surveys, we add random thermal noise and apply a simple foreground cleaning model, filtering out Fourier modes of the brightness temperature field with k ∥ < k min,∥ . Under current levels of thermal noise and foreground cleaning, typical of a Canadian Hydrogen Intensity Mapping Experiment (CHIME)-like survey, the ${\rm H\, \small {I}}$-galaxy cross-correlation signal in our simulations, using the kNN-field framework, is detectable at >30σ across r = [3, 12] h –1 Mpc. In contrast, the detectability of the standard two-point correlation function (2PCF) over the same scales depends strongly on the foreground filter: a sharp k ∥ filter can spuriously boost detection to 8σ due to position-space ringing, whereas a less sharp filter yields no detection. Nonetheless, we conclude that kNN-field cross-correlations are robustly detectable across a broad range of foreground filtering and thermal noise conditions, suggesting their potential for enhanced constraining power over 2PCFs.

79 ASTRONOMY AND ASTROPHYSICS↗

Theory of resonant x-ray scattering with ultrafast intense pulses

Here, we present a time-dependent Schrödinger equation approach within a nonrelativistic quantum electrodynamics framework to investigate resonant x-ray scattering driven by intense x-ray pulses. This method enables us to explore how coherent x-ray electron dynamics influence scattering signals from Ne + . We account for both resonance fluorescence and elastic scattering channels, while also considering competing photoionization and inner-shell decay processes. By computing the angular distribution and energy spectrum of scattered photons, we uncover interference effects between elastic scattering and resonance fluorescence pathways. Notably, this interference results in a small asymmetry in the energy spectrum. We discuss the experimental potential for detecting signatures of interference. Our findings demonstrate that the x-ray Rabi dynamics can be used to control scattering responses and provide insights into interference mechanisms and scattering efficiency in high-intensity x-ray regimes.

Venkatesh, Akilesh [Argonne National Laboratory (A↗

Nuclear mass radius and pressure in the Skyrme model

We compute the mass radius, scalar radius, tensor radius, baryon number radius, and mechanical radius of nuclei with baryon number B = 1 , 2 , 3 , 4 , 5 , 6 , 7 , 8 , 32, 108 in the Skyrme model. The relations between these radii and the nuclear gravitational form factors are investigated. We also compute the ‘pressure’ distribution and find that it is negative in the core region for all the nuclei with B > 1 . This suggests that the way mechanical stability is achieved in nuclei is qualitatively different than in the nucleon. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Engineering Privacy at the Edge: A Practical Guide to Differential Privacy in System Architectures

The rapid expansion of distributed and edge computing platforms—spanning autonomous vehicles, IoT sensors, and healthcare monitors—has heightened concerns about data privacy. Differential Privacy (DP) offers a rigorous mathematical framework to protect sensitive information while retaining analytical utility. This tutorial introduces the foundations of DP for both numerical and categorical datasets and extends the discussion to correlation-aware techniques tailored for structured and high-dimensional data. Hands-on demonstrations will begin with the PETINA (Privacy prEservaTIoN Algorithms) package for numerical data and continue with MIC-DP (Maximum Information Correlated Differential Privacy) for tabular data. Designed for researchers and practitioners in secure systems, embedded architectures, and AI accelerators, the tutorial emphasizes practical and scalable methods for integrating DP into real-world system designs.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Summary Report of the SOS26 Workshop held March 11-14, 2024

The SOS26 workshop was organized by Oak Ridge National Laboratory (ORNL) and held March 11-14, 2024, at Cocoa Beach, Florida. The SOS is a workshop organized annually, with a focus on distributed high performance computing (HPC). The technical program of the workshop was developed jointly by Sandia National Laboratories (SNL), ORNL and the Swiss National Supercomputing Center (CSCS). The 2024 SOS26 workshop theme was "Versatile HPC for the evolving and expanding needs of science" and had seven technical sessions covering HPC, data and machine learning (ML) topics. Each session consisted of four or five presentations, followed by a panel discussion. This report documents the workshop proceedings covering all the technical sessions.

97 MATHEMATICS AND COMPUTING↗

Size-Resolved Chemical Composition of Particles Collected Using STAC at the Ground Site During the SAIL Campaign in Gunnison, Colorado

Aerosol particles were collected using a four-stage Size and Time-resolved Aerosol Collector (STAC) during the SAIL field campaign. Each stage of STAC separates particles into distinct aerodynamic size fractions with 50% cut-off diameters: Stage A: 2.27 µm Stage B: 0.615 µm Stage C: 0.421 µm Stage D: 0.119 µm Each stage provides both size- and time-resolved sampling, enabling investigation of particle composition across different atmospheric regimes. Only a subset of samples was selected for analysis based on prevailing meteorological conditions (e.g., temperature, humidity, and air-mass influence) to capture representative aerosol types under distinct weather patterns. Collected substrates were first examined under Scanning Electron Microscopy (SEM) to evaluate particle loading, morphology, and spatial distribution. Subsequently, Computer-Controlled Scanning Electron Microscopy with Energy-Dispersive X-ray Spectroscopy (CCSEM/EDX) was performed to obtain size-resolved elemental composition of individual particles. A rule-based classification scheme was applied to categorize particles into major compositional groups (e.g., biological, carbonaceous, dust, sulfate, Na-rich, and mixed types). This dataset provides high-resolution morphological and chemical information on atmospheric particles collected during the SAIL campaign, offering insights into the influence of meteorology on aerosol composition and mixing state.

Size and Time-resolved Aerosol Collector↗

EUREICA: Efficient UltRa Endpoint IoT-enabled Coordinated Architecture

The electricity grid has evolved from a physical system to a cyber-physical system with digital devices that perform measurement, control, communication, computation, and actuation. The increased penetration of distributed energy resources (DERs) that include renewable generation, flexible loads, and storage provides extraordinary opportunities for improvements in efficiency and sustainability. However, they can introduce new vulnerabilities in the form of cyberattacks, which can cause significant challenges in ensuring grid resilience. The purpose of this project was to develop a framework ((Efficient, Ultra-REsilient, IoT-Coordinated Assets, or EUREICA)for achieving grid resilience through suitably coordinated assets including a network of Internet of Things (IoT) devices, and a local electricity market (LEM) to identify trustable assets and carry out this coordination. Situational Awareness (SA) of locally available DERs with the ability to inject power or reduce consumption is enabled by the market, together with a monitoring procedure for their trustability and commitment. Experiments conducted during this project demonstrated that, with this SA, a variety of cyberattacks can be mitigated using local trustable resources without stressing the bulk grid. The demonstrations were carried out using a variety of high-fidelity co-simulation platforms, real-time hardware-in-the-loop validation, and a utility-friendly simulator.

14 SOLAR ENERGY↗

A Review of Edge Computing Technology and Its Applications in Power Systems

Recent advancements in network-connected devices have led to a rapid increase in the deployment of smart devices and enhanced grid connectivity, resulting in a surge in data generation and expanded deployment to the edge of systems. Classic cloud computing infrastructures are increasingly challenged by the demands for large bandwidth, low latency, fast response speed, and strong security. Therefore, edge computing has emerged as a critical technology to address these challenges, gaining widespread adoption across various sectors. This paper introduces the advent and capabilities of edge computing, reviews its state-of-the-art architectural advancements, and explores its communication techniques. A comprehensive analysis of edge computing technologies is also presented. Furthermore, this paper highlights the transformative role of edge computing in various areas, particularly emphasizing its role in power systems. It summarizes edge computing applications in power systems that are oriented from the architectures, such as power system monitoring, smart meter management, data collection and analysis, resource management, etc. Additionally, the paper discusses the future opportunities of edge computing in enhancing power system applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗