Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Commuting embeddings for parallel strategies in non-local games

Non-local games provide a versatile framework for probing quantum correlations and for benchmarking the power of entanglement. In finite dimensions, the standard method for playing several games in parallel requires a tensor product of the local Hilbert spaces, which scales additively in the number of qubits. In this work, we show that this additive cost can be reduced by exploiting algebraic embeddings. We introduce two forms of compressions. First, when a referee selects one game from a finite collection of games at random, the game quantum strategy can be implemented using a maximally entangled state of dimension equal to the largest individual game, thereby eliminating the need for repeated state preparations. Second, we establish conditions under which several games can be played simultaneously in parallel on fewer qubits than the tensor product baseline. These conditions are expressed in terms of commuting embeddings of the game algebras. Moreover, we provide a constructive framework for building such embeddings. Using tools from Lie theory, we show that aligning the various game algebras into a common Cartan decomposition enables such a qubit reduction. Beyond the theoretical contribution, our framework casts NLGs as algebraic primitives for distributed and resource-constrained quantum computations and suggested NLGs as a comparable device-independent dimension witness.

Commuting embeddings↗

Experimental and Theoretical Evaluation of Feed Flow Collar Design for Shell Fed Hollow Fiber Membrane Modules

An experimental and theoretical study of module collar design is presented here. Hollow fiber membranes are prepared by dip coating a poly(vinylidene) (PVDF) support with a polydimethylsiloxane (PDMS) gutter layer and a Pebax 2533 selective layer. Fiber bundles with a well-defined fiber packing are prepared using a 3D printed module. A parallel fiber bundle consisting of 4-9 uniformly spaced fibers is created with printed tabs that align the fibers and create a tubesheet. The tabs are sealed within a printed case that possesses a series of external ports for gas introduction and removal. Uniquely, both port location and the use of a collar to assist fluid distribution in the shell can be varied for the same fiber bundle. Experimental measurements are compared to computational fluid dynamics (CFD) simulations. The experimental module design allows high-fidelity representation of the fiber bundle and module case in the simulations. Comparisons between experiment and simulation are in good agreement over a broad range of experimental conditions. The detrimental effect of having ports located too close, leading to stagnation regions, is captured as well as the beneficial effects of using a collar for shell-side fluid distribution around the fiber bundle. Such results help validate the use of CFD to develop high-performance module designs.

Tran, Thien↗

Synchrotron micro-computed tomography analysis of neutron-irradiated U-Mo fuel

The three-dimensional (3D) microstructure of neutron-irradiated uranium-10 wt.% molybdenum (U-10Mo) fuel with a burn-up of 9.8 × 10 21 fissions/cm 3 was characterized using a novel, multi-modal synchrotron micro-computed tomography approach combining propagation-based phase-contrast enhanced and absorption contrast techniques. The porosity development, porosity interconnectedness, swelling, composition, local thickness of the zirconium (Zr) diffusion barrier, and the influence of the fuel–cladding interaction on the local composition and pore morphology, were uniquely determined in 3D. Two cuboids were produced using a focused ion beam-scanning electron microscope at the Zr diffusion barrier–fuel interface and in the bulk fuel. The bulk fuel sample swelled by 53.3 [+9.7/−3.1]%, while the fuel near the Zr–fuel interface swelled by 63.3 [+14.7/−7.2]%. The average local thickness of the Zr diffusion barrier decreased by 53 %, compared to the expected pre-irradiated thickness. Four pore morphology regions were identified initiating parallel to the fuel–Zr interaction region: (1) an interaction layer of suppressed porosity, (2) a layer of elongated and interconnected porosity, (3) a transition zone of low porosity, and (4) a layer of unoriented porosity representative of the bulk fuel behavior. The increase in porosity near the diffusion barrier corresponded to a higher U concentration compared to that in the bulk fuel. The interconnected porosity in the fuel near the diffusion barrier was extensive and oriented parallel to the diffusion barrier, while the bulk fuel had more compact and isolated pore networks. The interaction layer, despite having suppressed porosity, was nearly 100 wt.% U. Porosity suppression at the diffusion barrier corresponds to the expected reduction in radiation-driven diffusion of Xe at the interface despite the anticipated increase in fission product nucleation originating from a higher U concentration. In conclusion, the novel 3D insights of the porosity, swelling, and compositional variations characterized herein can improve the fidelity of fuel performance codes for proliferation-resistant fuels for research and test reactors.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem

ChatHPC democratizes large language models for the high-performance computing (HPC) community by providing the infrastructure, ecosystem, and knowledge needed to apply modern generative AI technologies to rapidly create specific capabilities for critical HPC components while using relatively modest computational resources. Our divide-and-conquer approach focuses on creating a collection of reliable, highly specialized, and optimized AI assistants for HPC based on the cost-effective and fast Code Llama fine-tuning processes and expert supervision. We target major components of the HPC software stack, including programming models, runtimes, I/O, tooling, and math libraries. Thanks to AI, ChatHPC provides a more productive HPC ecosystem by boosting important tasks related to portability, parallelization, optimization, scalability, and instrumentation, among others. With relatively small datasets (on the order of KB), the AI assistants, which are created in a few minutes by using one node with two NVIDIA H100 GPUs and the ChatHPC library, can create new capabilities with Meta’s 7-billion parameter Code Llama base model to produce high-quality software with a level of trustworthiness of up to 90% higher than the 1.8-trillion parameter OpenAI ChatGPT-4o model for critical programming tasks in the HPC software stack.

Young, Aaron [ORNL] (ORCID:0000000254484667)↗

A novel approach to increase accuracy in remotely sensed evapotranspiration through basin water balance and flux tower constraints

Remote sensing-derived evapotranspiration (RSET) products capture the spatiotemporal variations of evapotranspiration (ET) from field to basin scales with unprecedented details. However, their accuracy varies across RSET estimation methods and diverse hydroclimate regions. While ET modeling efforts to account for biophysical processes and controlling parameters have made good progress in recent years, a parallel approach of integrating in-situ ET with RSET could reduce biases in RSET products. Basin water balance ET (WBET) and flux tower ET are widely applied to evaluate RSET accuracy, yet such ET measurements are rarely used for RSET bias corrections, especially for large area applications. To address this issue, we propose a novel approach: the water balance equivalence (WABE) method, which generates spatially continuous WBET for correcting biases in RSET products. The WABE method computes synthetic WBET by integrating observed WBET and flux tower-derived FLUXCOM ET, which fills the spatial gaps of observed WBET and generates a spatially continuous WBET dataset. Synthetic WBET (2002–2015 annual average) of eight-digit hydrologic unit code (HUC8) basins across the conterminous United States (CONUS), constituting 44 % (887 out of 2035 basins) of CONUS basins, was determined within 2.0 % (RMSE = 12 %) of observed WBET at CONUS and between 1–12 % (RMSE = 3–33 %) across 18 regions in CONUS. With WABE-based bias corrections, the overall annual bias of RSET decreased from 10 % (RMSE = 34 %) to 6 % (RMSE = 26 %) across 37 flux tower sites. The WABE method offers a new approach for RSET accuracy improvement and shows great promise for large area implementations with a potential to yield substantial benefits for building accurate basin water budgets and water management decisions.

Khand, Kul↗

A Full-Stack Exploration of Language-Based Parallelism in Fortran 2023

This poster explores native parallel features in Fortran 2023 through the lens of supporting applications with libraries, compilers, and parallel runtimes. The language revision informally named Fortran 2008 introduced parallelism in the form of Single Program Multiple Data (SPMD) execution with two broad feature sets: (1) loop-level parallelism via do concurrent and (2) a Partitioned Global Address Space (PGAS) comprised of distributed “coarray” data structures. Fortran’s native parallelism has demonstrated high performance [1] and reduced the burden of inserting what sometimes amounts to more directives than code. Several compilers support both feature sets, typically by translating do concurrent into serial do loops annotated by parallel directives and by translating SPMD/PGAS features into direct calls to a communication library. Our research focuses primarily on two questions: (1) can the compiler’s parallel runtime library be developed in the language being compiled (Fortran) and (2) can we define an interface to the runtime that liberates compilers from being hardwired to one runtime and vice versa. We are answering these questions by developing the Parallel Runtime Interface for Fortran (PRIF) [2] and the Co-Array Fortran Framework of Efficient Interfaces to Network Environments (Caffeine) [3]. Caffeine is initially targeting adoption by LLVM Flang, a new open-source Fortran compiler developed by a broad community in industry, academia, and government labs. We are also exploring the use of these features in Inference-Engine, a deep learning library designed to facilitate neural network training and inference for high-performance computing applications written in modern Fortran.

Rasmussen, Katherine↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Evidence of topological charge polarization at active-passive interfaces in acoustically powered active liquid crystals

Theoretical and computational studies predict topological charge separation at active-passive interfaces in active nematics, but reliable experimental validation has been lacking. We utilize a synthetic acoustically energized liquid crystal to experimentally investigate the dynamics of topological defects across an active-passive interface. The interference pattern induced by the acoustic wave inside the experimental cell creates a spatial distribution of activity, resulting in the formation of active-passive interfaces with the liquid-crystalline director field aligned parallel to those interfaces. The activity gradient drives the reorientation of positively charged topological defects along its direction, causing them to migrate into the passive zone and form a layer with net topological charge. That layer “discharges” upon cessation of activity through annihilation of positive and negative topological charges, resembling the discharge of a capacitor.

Active defects↗

SERAPH: Dark Matter Searches with SRF Cavities and Transmon Qubits

The Superconducting Quantum Materials and Systems Center, led by Fermi National Accelerator Laboratory, is one of five research centers funded by the U.S. Department of Energy as part of a national initiative to develop and deploy the world s most powerful quantum computers and sensors. SQMS will also apply the same technologies used for quantum computing, such as SRF cavities and superconducting qubits, to search for fundamental physics. This presentation will focus on the SERAPH experiment, a family of superconducting haloscopes being developed by SQMS to search for wavelike dark matter like axions and dark photons. In this presentation, I will focus on the progress of the current phase of SERAPH, which will search dark photon dark matter using a widely-tunable SRF cavity (4-7 GHz) with Q>10^8. In parallel, SQMS has recently developed superconducting transmon qubits with leading coherence times. I will report new results for SQMS dark matter searches implementing these qubits to subvert the Standard Quantum Limit noise.

79 ASTRONOMY AND ASTROPHYSICS↗

Ensemble Simulation Techniques and Fast Randomized Algorithms

The major goals of the project were to develop and analyze new ensemble simulation techniques, including trajectory stratification and preconditioned MCMC techniques, as well as develop fast numerical linear algebra techniques closely related to ensemble simulation ideas. The trajectory stratification techniques involve simulating in parallel short trajectory fragments of a Markov process confined to a specific region of space‐time and then patching together the statistics gathered to assemble estimates of very general dynamical properties. We have also developed this approach for rare event simulation and extended the techniques to applications requiring a more general framework (such as electronic structure calculations). The preconditioned MCMC techniques involve simulating multiple Markov chains in parallel and then using information from the ensemble to speed the mixing of each individual chain. The fast randomized linear algebra methods are motivated by the diffusion Monte Carlo technique, but are applicable to finding the dominant eigenvalue of (almost) general matrices. For most non‐negative matrices, the schemes result in an error (compared to the power method) that is constant in the dimension of the problem. For more general matrices, we see a very clear sublinear cost trend in computational tests.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja↗

GASNet-EX Specification Collection (Rev. 2024.5.0)

GASNet-EX is a portable, open-source, high-performance communication library designed to efficiently support the networking requirements of PGAS runtime systems and other alternative models in emerging exascale systems. It provides network-independent, high-performance communication primitives including Remote Memory Access (RMA) and Active Messages (AM). GASNet-EX is an evolution of the popular GASNet communication system, building upon over 20 years of lessons learned, and the primary goals are high performance, interface portability, and expressiveness. The library has been used to implement parallel programming models and libraries such as UPC, UPC++, Fortran coarrays, Legion, Chapel, and many others. This anthology collects together the four separate volumes that currently comprise the GASNet-EX specification, as of the 2024.5.0 release of GASNet-EX.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

GPU-Accelerated Solution of the Bethe–Salpeter Equation for Large and Heterogeneous Systems

We present a massively parallel GPU-accelerated implementation of the Bethe–Salpeter equation (BSE) for the calculation of the vertical excitation energies (VEEs) and optical absorption spectra of condensed and molecular systems, starting from single-particle eigenvalues and eigenvectors obtained with density functional theory. The algorithms adopted here circumvent the slowly converging sums over empty and occupied states and the inversion of large dielectric matrices through a density matrix perturbation theory approach and a low-rank decomposition of the screened Coulomb interaction, respectively. Further computational savings are achieved by exploiting the nearsightedness of the density matrix of semiconductors and insulators to reduce the number of screened Coulomb integrals. We scale our calculations to thousands of GPUs with a hierarchical loop and data distribution strategy. The efficacy of our method is demonstrated by computing the VEEs of several spin defects in wide-band-gap materials, showing that supercells with up to 1000 atoms are necessary to obtain converged results. We discuss the validity of the common approximation that solves the BSE with truncated sums over empty and occupied states. In conclusion, we then apply our GW-BSE implementation to a diamond lattice with 1727 atoms to study the symmetry breaking of triplet states caused by the interaction of a point defect with an extended line defect.

Absorption spectra↗

Contextual subspace variational quantum eigensolver calculation of the dissociation curve of molecular nitrogen on a superconducting quantum computer

Abstract We present an experimental demonstration of the Contextual Subspace Variational Quantum Eigensolver on superconducting hardware. Calculating the potential energy curve of molecular nitrogen proves challenging for many conventional quantum chemistry techniques, since static correlation dominates in the dissociation limit. Our quantum simulations retain good agreement with the Full Configuration Interaction energy, outperforming all benchmarked single-reference wavefunction techniques in capturing the bond-breaking appropriately. Moreover, our methodology is competitive with multiconfigurational approaches but at a saving of quantum resource, meaning larger active spaces can be treated for a fixed qubit allowance. To achieve this result, we deploy an error mitigation/suppression strategy comprised of Dynamical Decoupling, Measurement-Error Mitigation and Zero-Noise Extrapolation. Circuit parallelization also provides passive noise-averaging and improves the effective shot yield to reduce the measurement overhead. Furthermore, we introduce a modified adaptive ansatz construction algorithm that incorporates hardware awareness into our variational circuits, minimizing the transpilation cost for the target qubit topology.

Physics↗

Self-Consistent Relativistic Electron Scattering using the Sherlock Scattering Model for X-ray Diagnostics

We present on a new, self-consistent, arbitrary-temperature Romberg integration scheme for modeling electron scattering in materials in a LANL Lagrangian Shock Hydro (LSH) code. Electron beam-target interactions are fundamental to a wide range of scientific and technological applications. When high-energy electron beams hit their target, they may scatter, deposit energy, or ionize the source. These processes govern the behavior and outcomes in nanotechnology manufacturing, electron microscopy, and modern X-ray diagnostics. Simulating these interactions is essential for interpreting experimental results, predicting material responses, and designing efficient tools and experiments. At Los Alamos, this is done using a LSH code, which is a multi-dimension, multi-material, massively parallel, multi-physics code used to simulate applications from asteroid impacts to electron beam interactions. By effectively and efficiently modeling the way that electrons scatter from the beam we can bolster these simulations and more accurately predict experimental outcomes. The model currently implemented in the LSH of interest is based on work by Papp and does not self-consistently preserve momentum in the slightly relativistic regime; here we adopt a model proposed by Braams and Karney and implement a Romberg integration scheme to compute the diffusion tensor. In this paper we will provide background on the Braams-Karney diffusion tensor as well as the Romberg integration scheme we employed to numerically solve for it. We will show that our integration scheme is accurate in solving for the set of scalar potentials used to re-express the diffusion tensor in differential form, and in solving for the diffusion coefficients in the larger LSH code. By using this diffusion tensor rather than the existing Papp one, and numerically integrating it with a Romberg method, we produce much more accurate, self-consistent results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Alfalfa Virtual Building Service: Software Engineering Best Practices Applied to Runtime Interaction with Building Energy Models

Buildings are active participants in increasingly complex energy systems. Building Energy Modeling (BEM) has a key role to play in planning and de-risking an equitable energy transition, with BEM-backed "virtual buildings" critical path for diverse applications that include workforce training tools, Hardware-in-the-Loop (HIL) experimentation to study equipment performance under a range of conditions, Control-Hardware-in-the-Loop (CHIL) experimentation to de-risk commercial control implementations at equipment through grid orchestration levels, and integration of dynamic load profiles into grid modeling tools for energy system experimentation at the urban scale. Modeling requirements vary across these applications, but many software engineering tasks do not. The Alfalfa Virtual Building Service (AVBS, see https://github.com/NREL/alfalfa/wiki) is an open-source web service that solves these common tasks robustly in one place, providing a foundational platform for power users to bootstrap their own applications. AVBS abstracts the specifics of runtime interaction with OpenStudio, Modelica, and Spawn of EnergyPlus models behind a unified REST API. Additionally, AVBS provides resources for cloud deployment and scaling to 100s of parallel simulations, a growing library of modular Operational Technology (OT) integrations for emulation of real-world interfaces, and scripts to automate the population of communities of virtual buildings from URBANopt, ResStock and ComStock.

building automation↗