Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Accelerating Application Bulk Synchronous Writes in HPC Environments

High-bandwidth storage tiers are becoming more common for their capability to absorb high-rate, bursty I/Os. Notably, the designs of these fast storage tiers differ from system to system. The variation of these layers and non-uniform methods of access can pose chal- lenges for applications seeking to run at multiple HPC facilities. Therefore, in this work, we present Spectral, a rapid-output ab- straction library to accelerate application, bulk-synchronous writes on HPC systems. We design Spectral to enable applications to use high-bandwidth storage, such as node-local storage and dis- tributed, write-caches (e.g., burst buffers) transparently without requiring modifications to the application or file system source code. The key idea is to allow applications to spend most of the time performing productive work and to not require any source code changes for maximum portability on different HPC archi- tectures. Spectral internally re-routes write-only files through available, high-performance I/O resources before ultimately mi- grating them to the shared global parallel file system. For instance, on Summit, Spectral transparently places application outputs on node-local storage and then utilizes asynchronous migration to the center-wide GPFS file system. We evaluate Spectral on the Summit HPC system (1024 nodes) using the IOR benchmark and real scientific applications. Spectral shows linear performance scaling, improving application write performance by over an order of magnitude when compared to GPFS.

Khan, Awais↗

Trigonometric continuous-variable gates and hybrid quantum simulations of the sine-Gordon model

Hybrid qubit-qumode quantum computing platforms provide a natural setting for simulating interacting bosonic quantum field theories. However, existing continuous-variable gate constructions rely predominantly on polynomial functions of canonical quadratures. In this work, we introduce a complementary universality paradigm based on trigonometric continuous-variable gates, which enable a Fourier-like representation of bosonic operators and are particularly well suited for periodic and non-perturbative interactions. We present an ancilla-based framework for implementing trigonometric gates with arguments given by arbitrary Hermitian functions of qumode quadratures. The protocol yields unitary gates deterministically, and non-unitary gates through probabilistic post-selection. As a concrete application, we develop a hybrid qubit-qumode quantum simulation of the lattice sine-Gordon model. Using these gates, we prepare ground states via quantum imaginary-time evolution, simulate real-time dynamics, compute time-dependent vertex two-point correlation functions, and extract quantum kink profiles under topological boundary conditions. Our results demonstrate that trigonometric continuous-variable gates provide a physically natural framework for simulating interacting field theories on near-term hybrid quantum hardware, while establishing a parallel route to universality beyond polynomial gate constructions. We expect that the trigonometric gates introduced here to find broader applications, including quantum simulations of condensed matter systems, quantum chemistry, and biological models.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

Mineralogical evidence for hydrothermal alteration of Bennu samples

Samples of asteroid (101955) Bennu delivered by the OSIRIS-REx mission offer the opportunity to study pristine planetary materials unchanged by exposure to the terrestrial environment. Here we use a combination of X-ray diffraction and various electron microscopy techniques to explore the detailed mineralogy of Bennu samples and determine the alteration history of the planetesimal protolith from which they originated. The samples consist largely of hydrated sheet-silicate minerals, namely nanoscale serpentine and saponite of varied grain size, which are decorated with micro- to nanoscale Fe-sulfides, magnetite and carbonates. We observe sheet silicates parallel and normal to sulfide surfaces and as inclusions in sulfides; sulfur-rich veins transecting the sheet-silicate matrix; zoned carbonates and phosphates and sulfide and magnetite grains exhibiting embayment. The mineralogical evidence indicates alteration of accreted minerals by a fluid that evolved with time, leading to etching, dissolution and reprecipitation. Sulfide compositions indicate alteration at ~25 °C, similar to conditions inferred for asteroid (162173) Ryugu and Ivuna-type (CI) chondrite meteorites. The fluid probably evolved from neutral to alkaline, culminating with the precipitation of highly soluble salts. We conclude that Bennu’s protolith comprised mainly nanometre to micrometre silicates, with fewer chondrules and calcium–aluminium-rich inclusions than those of most chondrite groups.

Zega, T J↗

Virtual Cable Impedance based Load Sharing in a Microgrid for Parallel Connected Grid Forming Converters

This paper presents a novel approach to power sharing between direct connected grid-forming converters, utilizing virtual cable impedance and droop-based outer loop control. To enhance stability, resistive droop is implemented, while virtual cable impedance with non-zero resistive and inductive components ensures improved power sharing. The inner loop controller employs a Lyapunov energy function to achieve superior dynamic performance. The proposed control architecture is validated through comprehensive modeling and real-time processor-in-the-loop simulations, demonstrating its robustness and efficiency under various operating conditions. The results highlight the potential of this control strategy to improve the reliability and efficiency of renewable energy systems. Additionally, a comparative analysis with traditional methods underscores the advantages of the proposed approach in terms of stability and performance. The proposed control architecture offers a scalable and flexible solution for grid-forming converters, enabling seamless integration of renewable energy sources. Its robustness and adaptability make it an attractive solution for real-world applications. Furthermore, the approach can be extended to other power electronic systems, enhancing overall system performance and efficiency. By providing a reliable and efficient control strategy, this paper contributes to the advancement of renewable energy systems and their adoption in the energy sector. The proposed control strategy has far-reaching implications for the widespread adoption of renewable energy sources, enabling a more sustainable and efficient energy future. The overall system is modeled in MATLAB/Simulink and PLECS software domain.

grid forming converters (GFM)↗

Experimental observations of bifurcated power decay lengths in the near Scrape-Off Layer of ST40 High Field Spherical Tokamak

The scrape-off layer parallel heat flux decay lengths measured at ST40, a high field, low aspect ratio spherical tokamak, have been observed to bifurcate into two groups. The wide group follows established H-mode scalings (ranging between 2 to 8 mm) while the narrow group falls up to 10 times below these scalings (between 0.2 and 0.8 mm), being comparable to the ion total Larmor radius rather than the ion poloidal Larmor radius. The heat flux profiles of the latter group can only be described by a multi-exponential function, rather than the single exponential function convoluted with a Gaussian. The onset of the narrow scrape-off layer width is observed to be associated with suppressed magnetic fluctuations, suggesting reduced electromagnetic turbulence levels in the SOL.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Snow ALbedo eVOlution (SALVO) Campaign Spectral Albedo and Related Measurements from April - June, 2024 in Utqiagivk, AK

A field-portable spectroradiometer, referred to herein as an ‘ASD’, was used to make spatially-distributed spectral albedo (350 – 2500 nm) measurements on tundra and sea ice surfaces. The ASD detector is carried in a backpack and controlled via a computer mounted on the front of the operator (see Figure 1). The ASD measures the spectral irradiance from a fiber optic cable that is routed from the backpack to a custom, gooseneck cosine collector mounted on the end of a 1-m long boom (Grenfell and Perovich, 2008). The boom was held at hip height (approximately 1 m) and had an integrated bubble level for levelling. To make an albedo measurement, first the operator collect an incident (down-welling) irradiance, followed by a reflected (up-welling) measurement. The time between incident and reflected measurements was typically between 11 and 26 seconds (interquartile range). For each measurement, 10 spectra are averaged together. Albedo is calculated as the ratio of the reflected to incident measurement, which obviates the need for absolute radiometric calibration. Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1 m south of the line. While the ASD operator was making measurements, an assistant kept notes on the scan number associated with each measurement, the surface type (see below), and collected photos of each measurement (see companion oblique photos data archive). Measurements were made within 3 hours of solar noon.

ASD Spectroradiometer↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

Parallel expansion of a fuel pellet plasmoid

The problem of the assimilation of a cryogenic fuel pellet injected into a hot plasma is considered. Due to the transparency to ambient particles of the plasmoid, the localised region of high-density plasma created by ionisation of the ablated pellet material, electrons reach a ‘quasiequilibrium’ (QE) state which is characterised by a steady-state on the fastest collisional time scale. The simplified electron kinetic equation of the QE state is solved. Taking a velocity moment of the higher-order electron kinetic equation, which is valid on the expansion time scale, permits a fluid closure, yielding an evolution equation for the macroscopic parameters describing the QE distribution function. In contrast to the Braginskii equations, the closure does not require that electrons have a short mean free path compared with the size of density perturbations, and permits an anisotropic and highly non-Maxwellian distribution function. As the QE distribution function accounts for both trapped and passing electrons, the self-consistent electric potential that causes the expansion can be properly described, in contrast to earlier models of pellet plasmoid expansion with an unbounded potential. The plasmoid expansion is simulated using both a Vlasov model and a cold-fluid model for the ions. During the expansion plasmoid ions and electrons obtain nearly equal amounts of energy; as hot ambient electrons provide this energy in the form of collisional heating of plasmoid electrons, the expansion of a pellet plasmoid is expected to be a potent mechanism for the transfer of energy from electrons to ions on a time scale shorter than that of ion–electron thermalisation.

Physics↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Dynamic magneto-chiral instability in photoexcited tellurium

In systems of charged chiral fermions out of equilibrium, an electric current parallel to a magnetic field can generate a dynamic instability that amplifies electromagnetic waves. Whether this mechanism also operates in chiral solid-state systems has remained uncertain. Here we observe signatures of a dynamic magneto-chiral instability in elemental tellurium, a structurally chiral crystal, using time-domain terahertz emission spectroscopy. Under transient photoexcitation in a moderate magnetic field, we observe terahertz radiation with coherent modes that grow in amplitude over time. We present a theoretical model that describes this behaviour based on a dynamic instability of electromagnetic waves interacting with infrared-active oscillators of acceptor states in tellurium, giving rise to an amplifying polariton. These results demonstrate that magneto-chiral instabilities can emerge in solid-state systems and establish a mechanism for terahertz-wave amplification in chiral materials.

42 ENGINEERING↗

Similarity for downscaled kinetic simulations of electrostatic plasmas: Reconciling the large system size with small Debye length

A simple similarity has been proposed for kinetic (e.g., particle-in-cell) simulations of plasma transport that can effectively address the long-standing challenge of reconciling the tiny Debye length with the vast system size. This applies to both transport in unmagnetized plasma and parallel transport in magnetized plasmas, where the characteristics length scales are given by the Debye length, collisional mean free paths, and the system or gradient lengths. The controlled scaled variables are the configuration space, x/L, and an artificial Coulomb Logarithm, L ln Λ, for collisions, while the scaled time, t/L, and electric field, LE, are automatic outcomes. The similarity properties are examined, demonstrating that the macroscopic transport physics is preserved through a similarity transformation while keeping the microscopic physics at its original scale of Debye length. To showcase the utility of this approach, two examples of 1D plasma transport problems were simulated using the VPIC code: the plasma thermal quench in tokamaks [Li et al., Nuclear Fusion 63, 066030 (2023)] and the plasma sheath in the high-recycling regime [Li et al., Physics of Plasmas 30, 063505 (2023)].

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Assessment of ESM Readiness Level for Exascale HPC

Advancement of Earth System Models (ESMs) is becoming increasingly challenging due to a confluence of factors including increasing model complexity – to more fully represent the earth system, increasing spatial resolution - to achieve higher accuracy by resolving fine-scale dynamical to physical, biological, and chemical processes and their interaction, increasing ensemble size - to more accurately represent predictive uncertainty, and increased computing requirements – to enable more accurate and timely weather predictions and climate projections for societal benefit. The belief by many that computing will take care of itself is no longer valid given the disruptive changes in HPC that are driving up the cost of computing, increasing the difficulty of using emerging HPC effectively, and exposing limits in parallelism, portability and scalability of the ESM applications themselves.

54 ENVIRONMENTAL SCIENCES↗

Tailoring microstructures with mild magnetic-field processing: A case study of CuNiFe alloys

Combined experimental and computational investigations of the CuNiFe spinodal system confirm that application of a mild magnetic field during thermal treatment alters elemental redistribution and the resulting microstructure, relative to that obtained from zero-field annealing. Spinodal decomposition of a Cu 40 Ni 42 Fe 18 alloy was initiated during thermal treatment at 773 K, conducted either under zero field or modest (60 mT) magnetic f ield conditions for up to 200 h. Periodic (~10 nm) chemical modulations into Cu-rich and NiFe-rich regions were observed under both conditions, with the amplitude and wavelength of the segregated regions increasing with treatment time. However, magnetic field annealing resulted in a more than twofold increase in the amplitude of elemental modulations relative to zero-field conditions – consistent with enhanced diffusional f luxes during spinodal decomposition – while the modulation wavelength remained largely unaffected. These microstructural differences are reflected in various extrinsic magnetic properties. In parallel, first-principles DFT calculations indicate that long-range ferromagnetic order, as induced by an applied magnetic field, substantially alters the strength and nature of atomic interactions, enhancing the thermodynamic instability of the CuNiFe solid solution. Collectively, these results suggest that incorporating a mild (millitesla-level) magnetic field – distinct from the strong (tesla-level) fields commonly used in prior studies – during thermal processing has the potential to deliver enhanced control of microstructures for targeted engineering outcomes.

36 MATERIALS SCIENCE↗

Testing and Analysis of Grid Forming Inverter Control for Achieving Resilient and Economic Operation of an Islanded Microgrid

This investigation examines the feasibility of operating a battery energy storage system (BESS) in parallel with synchronous generation by using grid forming (GFM) control in order to achieve frequency control objectives while mitigating increases to operating costs in the context of an islanded microgrid. The BESS GFM control system, which is based on conventional droop techniques, is modeled along with the overall microgrid using the Real Time Digital Simulator (RTDS) to allow for integration of genset controller hardware. A series of simulations are performed to test the voltage and frequency regulation capability of the BESS control system when the primary frequency regulating genset is tripped offline. The results of the simulations suggest that the GFM control scheme will successfully maintain frequency and voltage stability, which will enable operation without a back-up genset while not compromising the microgrid resiliency to contingencies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Integrating Flow Imaging Analysis and Single-Particle ICP-TOFMS for Comprehensive Micro- and Nanoplastic Characterization

Flow imaging analysis (FIA), provides composition-agnostic morphological characterization. These measurements of particle size and shape are valuable to mass-based analysis, such as single particle inductively coupled plasma time-of-flight mass spectrometry (sp-ICP-TOFMS), which provides quantitative data on elements within particles. Using these two methods together enables informed use of geometric assumptions required by sp-ICP-TOFMS, as particle mass is typically converted to a particle diameter using assumed-spherical geometry. To validate this concept, parallel measurements to determine particle diameters were performed by FIA and sp-ICP-TOFMS on four particle suspensions: 300 nm polystyrene Eu-doped nanoparticles, 1 μm Fe-rich beads, 3 μm four element calibration polystyrene beads and 5 μm polystyrene beads. The Fe-particles obtained the highest percent difference from the manufacturer’s nominal diameter, as the mean diameter obtained by FIA was overestimated by 21% and sp-ICP-TOFMS underestimated the mean diameter by 20.7%. Two types of particles were selected to test the effect of varying the particle number concentrations (PNC) on sizing accuracy, and both methods accurately sized each particle population at the PNC expected. Single particle analysis of carbon has continued to be a popular research topic, with direct applications to environmental pollutants in terms of nano- and micro- plastics. Real-world plastic particles were studied, and FIA’s measured circularity values demonstrated that the particles deviated from spherical geometries, therefore sp-ICP-TOFMS data should be interpreted as mass-based rather than size-based. Combining these techniques enables improved interpretation of particle populations and evaluation of particle sizes.

Szakas, Sarah [ORNL] (ORCID:0000000241332197)↗

Spherical tokamak physics research in preparation for the operation of NSTX-U

The National Spherical Torus Experiment Upgrade (NSTX-U) is preparing to resume operation, representing a crucial step toward realizing compact, cost-effective fusion pilot plants. In advance of this, extensive modeling and data analysis have been conducted to advance the physics basis for low-aspect-ratio, high-performance plasma regimes, focusing on three core objectives: confinement and stability, power and particle handling, and steady-state operation. Significant progress has been made in understanding the electron temperature flattening in high-β plasmas, which is shown to be driven by a complex interplay of magnetohydrodynamic instabilities (e.g. non-resonant infernal modes), fast-ion-driven Alfvén eigenmodes, and electron and ion-scale micro-instabilities, particularly Kinetic Ballooning Modes (KBMs), whose destabilization is strongly dependent on parallel magnetic field fluctuations (δB ∥ ). Furthermore, a new gyrokinetic critical pedestal model was developed, accurately predicting pedestal structure by identifying KBMs as the primary stability limit, offering a critical constraint for future high-confinement scenarios. To address the challenge of high heat flux, novel liquid lithium plasma-facing components were modeled. The analysis confirmed that lithium vapor shielding is a self-regulating mechanism for heat mitigation, while also emphasizing that strong main ion parallel flow is essential to minimize core lithium contamination. Finally, progress toward steady-state operation was anchored by developing the required physics basis and control tools. This includes predictive modeling for reversed magnetic shear sustainment, demonstrating that magnetic island-induced bootstrap current reduction is negligible in STs, and advancing real-time control and disruption avoidance capabilities. The development of high-speed surrogate models (e.g. MMMNet) provides computationally efficient tools vital for non-inductive scenario optimization and integrated, low-disruptivity operations planned for NSTX-U.

NSTX-U↗