Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Adaptive Grid Redistribution for a 1D Model of Turbulence and Clouds

In global atmospheric models, resolving stratocumulus (Sc) in the vertical is computationally expensive. However, Sc appear only under special meteorological conditions. Therefore, there is motivation to refine the vertical grid levels adaptively. In order to facilitate the possibility of parallelization on graphical processing units, our grid adaptation method prescribes the number of vertical levels a priori. Then grid levels are relocated toward altitude ranges in need of refinement. Because the method relocates existing grid levels, rather than adding extra levels, there is a risk of creating regions with overly coarse grid spacing, that is, voids in the grid mesh. To prevent such voids from forming, a simple method is developed to impose a maximum grid spacing. To decide where to place enhanced resolution, the authors develop an empirical mesh refinement criterion. It refines grid spacing near the ground, near strong temperature gradients, and within clouds. Our grid adaptation method is implemented in a single-column model and evaluated on four test cases: decaying stratocumulus, developing shallow cumulus, a quasi-stationary stratocumulus deck, and the diurnal cycle of a dry boundary layer. In the stratocumulus cases, mesh refinement leads to improvements in both the time evolution of fields and their time averages. The other two cases show smaller differences.

Carstensen, Steffen [Univ. of Wisconsin, Milwaukee↗

Reconstruction of neutrino events in the Accelerator Neutrino Neutron Interaction Experiment. Part I

The Accelerator Neutrino Neutron Interaction Experiment (ANNIE) was designed to reconstruct neutrino events from the Fermilab Booster Neutrino Beam (BNB) with the parallel goals of measuring neutron production in interactions with oxygen and serving as a testbed for new technology. The ANNIE detector consists of a 26-ton water Cherenkov target tank instrumented with conventional photomultiplier tubes (PMTs), a downstream tracking muon spectrometer, and an upstream double wall of plastic scintillator to serve to veto charged particles incoming from neutrino events that occur upstream of the experimental setup. ANNIE has also deployed multiple Large-Area Picosecond PhotoDetectors (LAPPDs) and a test vessel of water-based liquid scintillator (WbLS). This paper describes the event reconstruction performance of the detector before implementation of these novel technologies, which will serve as a baseline against which their impact can be measured. That said, even the techniques used for event reconstruction using only the conventional PMT array and muon spectrometer are significantly different than those used in other water Cherenkov detectors due to the small size of ANNIE (which makes nanosecond-scale timing not as useful as in a large detector) and the availability of reconstruction information from the tracking muon spectrometer. We demonstrate that combining the information from these two elements into a single fit using only pattern recognition yields a muon vertex uncertainty of 60 cm, a directional uncertainty of 13.2 degrees, and energy reconstruction uncertainty of about 10% for BNB muon neutrino Charged Current Zero Pion (CC0π) events.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Enamel nanocrystal misorientation increased with meat-eating and agriculture

Enamel covers teeth, is the hardest tissue in the vertebrate body and has a complex multiscale structure from nanometres to millimetres. The structure comprises thin, long hydroxyapatite (Ca 5 (PO 4 ) 3 OH) nanocrystals, 50–70 nm wide, many micrometres long, parallel and bundled into approximately 5-µm-wide rods. The rods undulate and cross into a microscale ‘decussation pattern’ that toughens enamel by deflecting cracks. However, the crystallographic orientation of enamel nanocrystals is poorly understood. Here we show that the misorientation angle of adjacent nanocrystals varies markedly across 12 primate teeth spanning 9 species, 17.8 million years of evolution and diverse diets. Using a method called Polarization Enabled Large Input of Crystal Angles at the Nanoscale (PELICAN), we compare nanocrystals in the same (pre)molar locations and show that misorientation increases with food hardness in extant and fossil non-human apes and monkeys. We compare misorientation across three major dietary shifts in human evolution: the transition to meat-eating about 2.0–1.5 million years before present, to agriculture (about 12,000 years before present), and the Industrial Revolution (about 250 years before present). We show that over the past 1.6 million years, in the human lineage misorientation increased with time, especially when meat and stone-ground grains were introduced into human diets, but not with the Industrial Revolution. Thus, besides macro-changes, teeth adapted to dietary change at the nanoscale and crystallographically. This observation suggests that misorientation may contribute to enamel’s resilience; thus, bioinspired materials may consider small misorientation angles for added resilience.

biomaterials↗

Ion Cyclotron Heating in a Levitated Dipole Fusion Reactor

OpenStar Technologies is pursuing the levitated dipole (LD) as a highly modular, loosely-coupled system that leverages their expertise in high temperature superconductor (HTS) technology. The next generation experiment at OpenStar, Tahi (Ma¯ori for “first”), will demonstrate the generation and confinement of fast ions in a levitated dipole for the first time. Ion cyclotron range of frequency (ICRF) heating is a leading candidate for energetic ion formation in Tahi. A frequency in the 10 MHz range will be used for H minority heating or D majority heating with waves launched from an antenna located above the floating coil. Unlike a tokamak, where the targeted cyclotron resonance is typically a vertical path through the center of the plasma, in a dipole the resonance location follows a C-shaped path from the separatrix to the center of the plasma. The value of B also varies significantly within the confined plasma resulting in a large number of cyclotron harmonics present in the low field region. Furthermore, levitated dipoles contain a “first closed flux surface” surrounding the floating coil, in addition to the traditional separatrix /last closed flux surface. Simulation e ff orts using full-wave ICRF codes show that ICRF heating of a levitated dipole reactor is feasible using a pair of toroidal current straps phased to launch the appropriate parallel ( i.e. poloidal) refractive index.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Latent Twins

Over the past decade, scientific machine learning has transformed the development of mathematical and computational frameworks for analyzing, modeling, and predicting complex systems. From inverse problems to numerical partial differential equations (PDEs), dynamical systems, and model reduction, these advances have pushed the boundaries of what can be simulated. Yet they have often progressed in parallel, with representation learning and algorithmic solution methods evolving largely as separate pipelines. With Latent Twins, we propose a unifying mathematical framework that creates a hidden surrogate in latent space for the underlying equations. Whereas digital twins mirror physical systems in the digital world, Latent Twins mirror mathematical systems in a learned latent space governed by operators. Through this lens, classical modeling, inversion, model reduction, and operator approximation all emerge as special cases of a single principle. We establish the fundamental approximation properties of Latent Twins for both ordinary differential equations (ODEs) and PDEs and demonstrate the framework across three representative settings: (i) canonical ODEs, capturing diverse dynamical regimes; (ii) a PDE benchmark using the shallow-water equations, contrasting Latent Twin simulations with deep operator network and forecasts with a four-dimensional variational method baseline; and (iii) a challenging real-data geopotential reanalysis dataset, reconstructing and forecasting from sparse, noisy observations. Latent Twins provide a compact, interpretable surrogate for solution operators that evaluate across arbitrary time gaps in a single-shot, while remaining compatible with scientific pipelines such as assimilation, control, and uncertainty quantification. Looking forward, this framework offers scalable, theory-grounded surrogates that bridge data-driven representation learning and classical scientific modeling across disciplines.

Latent Twins↗

Thermal and kinematic properties of ejecta in SN1987A revealed by XRISM

We present an analysis of high-resolution spectra from the shock-heated plasmas in SN 1987A, based on an observation using the Resolve instrument onboard the X-Ray Imaging and Spectroscopy Mission (XRISM). The 1.7–10 keV Resolve spectra are accurately represented by a single-component, plane-parallel shock plasma model, with a temperature of $2.84_{-0.08}^{+0.09}$ keV and an ionization parameter of $2.64_{-0.45}^{+0.58}$ × $10^{11}\,\,{\rm s\,\, cm}^{-3}$. The Resolve spectra are also well reproduced by the 3D magneto-hydrodynamic simulation presented by Orlando et al. (2020, A&A, 636, A22) suggesting substantial contribution from the ejecta. The metal abundances obtained with Resolve align with the Large Magellanic Cloud value, indicating that the X-rays in 2024 originate from “non-metal-rich” shock-heated ejecta and the reverse shock has not reached the inner metal-rich region of ejecta. Doppler widths of the atomic lines from Si, S, and Fe correspond to velocities of 1500–1700 km s$^{-1}$, where the thermal broadening effects in this non-metal-rich plasma are negligible. Therefore, the line broadening seen in Resolve spectra is determined by the large bulk motion of ejecta. For reference, we determined a $90\%$ upper limit on non-thermal emission from a pulsar wind nebula at $4.3 \times 10^{-13}$ erg cm$^{-2}$ s$^{-1}$ in the 2–10 keV range, aligning with NuSTAR findings by Greco et al. (2022, ApJ, 931, 132). Additionally, we searched for the $^{44}$Sc K line feature and found a $1\sigma$ upper limit of $1.0 \times 10^{-6}$ photons cm$^{-2}$ s$^{-1}$, which translates to an initial $^{44}$Ti mass of approximately $2 \times 10^{-4}\, M_{\odot }$, consistent with previous X-ray to soft gamma-ray observations (Boggs et al. 2015, Science, 348, 670; Grebenev et al. 2012, Nature, 490, 373; Leising 2006, ApJ, 651, 1019).

ISM: supernova remnants↗

Modernizing the Legacy Fission Wire Measurement System for the Advanced Test Reactor-Critical Facility

Operational lifetime extensions of existing research reactors have emphasized the need for refurbishment, replacements, and upgrades to supporting equipment and instrumentation. The Advanced Test Reactor (ATR) at Idaho National Laboratory (INL), which entered service in 1967, has recently completed the sixth core internals change-out and has scheduled operations until at least 2040. Reactor maintenance and operational risk management is critically important in the research reactor community, however supporting measurement systems sometimes get overlooked when maintenance is planned. The Fission Wire Measurement System (FWMS) is a custom measurement system designed in the 1960s to measure the beta-particle activity of irradiated uranium-aluminum fission wires. This measurement is conducted to determine the fission rate profile of the Advanced Reactor Test Critical (ATR-C) facility. The ATR-C is an open-pool, low-power test reactor that was purpose driven to resemble ATR and is used to qualify experiment configurations and verify core models prior to full-power experiment irradiations in ATR. A power distribution measurement in ATR-C uses uranium-aluminum wires that are distributed throughout the ATR-C core to validate simulation and modeling results. These measurements require 340 to 1500 wires to be irradiated and measured within a 12-hour window. The activity of the wires is measured in the required time with the FWMS, which was put into service in 1965 at the Radiation Measurements Laboratory (RML). The system consists of 4 measurement channels and one reference channel, each with a 2-pi proportional gas flow detector and the measurement channels each have an automated sample changer. This legacy system is crucial to the continued operations of ATR and has undergone some minor hardware upgrades since 1965, however the system presently relies on custom control boards, custom gas ion chambers, analog amplifiers/discriminators, and a user interface (UI) for the system written in outdated code. Much of the equipment and software is custom with no commercial replacements or support and limited documentation. The existing control software requires an operating system that is no longer supported, creating more vulnerabilities to continued operations. A project is underway with a third-party vendor to design, build, and document a new control and data acquisition system (CDAS) for the FWMS. The new upgrade will replace the control system, computer, UI, sample changer motors, and main power supply while maintaining the interface with existing detector hardware. The upgraded system will be operated in parallel with the current hardware and software to conduct validation testing. This equipment upgrade demonstrates the commitment at ATR to ensuring successful operations and potential future research reactors at INL.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Short-Wave Infrared Upconverting Nanoparticles

Optical technologies enable real-time, noninvasive analysis of complex systems but are limited to discrete regions of the optical spectrum. While wavelengths in the short-wave infrared (SWIR) window (typically, 1700-3000 nm) should enable deep subsurface penetration and reduced photodamage, there are few luminescent probes that can be excited in this region. Here, we report the discovery of lanthanide-based upconverting nanoparticles (UCNPs) that efficiently convert 1740 or 1950 nm excitation to wavelengths compatible with conventional silicon detectors. Screening of Ln3+ ion combinations by differential rate equation modeling identifies Ho3+/Tm3+ or Tm3+ dopants with strong visible or NIR-I emission following SWIR excitation. Experimental upconverted photoluminescence excitation (U-PLE) spectra find that 10% Tm3+-doped NaYF4 core/shell UCNPs have the strongest 800 nm emission from SWIR wavelengths, while UCNPs with an added 2% or 10% Ho3+ show the strongest red emission when excited at 1740 or 1950 nm. Mechanistic modeling shows that addition of a low percentage of Ho3+ to Tm3+-doped UCNPs shifts their emission from 800 to 652 nm by acting as a hub of efficient SWIR energy acceptance and redistribution up to visible emission manifolds. Parallel experimental and computational analysis shows rate equation models are able to predict compositions for specific wavelengths of both excitation and emission. These SWIR-responsive probes open a new IR bioimaging window, and are responsive at wavelengths important for vision technologies.

Qi, Xiao↗

ACHILLES-GENIE Interface for Neutrino Simulations, ADRIANO2 Tile Prototype for High-Granularity Dual-Readout Calorimetry

ICARUS (Imaging Cosmic And Rare Underground Signals) is a liquid argon time projection chamber (LArTPC) detector that pursues the sterile neutrino, which relies on accurate simulations of neutrino-argon interactions. REDTOP (Rare Eta Decays To Observe new Physics) is a proposed low-energy, high-intensity meson factory designed to explore rare $\eta$/$\eta'$ meson decays and probe physics beyond the Standard Model. As a next-generation experiment, this requires both accurate simulations and innovative detector technologies. This project contributes to both ICARUS, from a simulation perspective, and REDTOP, from both a simulation and detection perspective, through the event generation of lepton-nucleon interactions and the physical enhancement of the calorimeter technology within the REDTOP detector. We developed an interface between ACHILLES (A CHIcago Land Lepton Event Simulator), a theory-driven lepton-level event generator, and GENIE, a robust event generator framework used for neutrino physics. By incorporating the precise theoretical cross-section calculations of ACHILLES into the experimental realism of GENIE, the interface allows for improved accuracy of neutrino-nucleon simulations, which can be adapted for the proton beam specifications of the REDTOP meson factory as well as for the ICARUS experiment. In parallel, we developed an improved prototype for the ADRIANO2 (A Dual Readout Integrally Active Non-segmented Option) dual-readout calorimeter tiles for the REDTOP detector. To improve the efficiency of the lead-glass tiles trapping Cherenkov light for energy reconstruction and particle identification, we optimized the application of a highly reflective coating. Through viscosity and thickness control, masking, and a custom spray technique, we refined the coating process to reduce surface defects and improve light yield. Together, these efforts strengthen the ICARUS neutrino program and REDTOP's capability of detecting rare decay events.

Visser, Erin [Michigan State U.]↗

Enhancing ICARUS and REDTOP Software and Hardware: Event Generator Interface Development and Calorimeter Tile Prototype

ICARUS (Imaging Cosmic And Rare Underground Signals) is a liquid argon time projection chamber (LArTPC) detector that pursues the sterile neutrino, which relies on accurate simulations of neutrino-argon interactions. REDTOP (Rare Eta Decays To Observe new Physics) is a proposed low-energy, high-intensity meson factory designed to explore rare $\eta$/$\eta'$ meson decays and probe physics beyond the Standard Model. As a next-generation experiment, this requires both accurate simulations and innovative detector technologies. This project contributes to both ICARUS, from a simulation perspective, and REDTOP, from both a simulation and detection perspective, through the event generation of lepton-nucleon interactions and the physical enhancement of the calorimeter technology within the REDTOP detector. We developed an interface between ACHILLES (A CHIcago Land Lepton Event Simulator), a theory-driven lepton-level event generator, and GENIE, a robust event generator framework used for neutrino physics. By incorporating the precise theoretical cross-section calculations of ACHILLES into the experimental realism of GENIE, the interface allows for improved accuracy of neutrino-nucleon simulations, which can be adapted for the proton beam specifications of the REDTOP meson factory as well as for the ICARUS experiment. In parallel, we developed an improved prototype for the ADRIANO2 (A Dual Readout Integrally Active Non-segmented Option) dual-readout calorimeter tiles for the REDTOP detector. To improve the efficiency of the lead-glass tiles trapping Cherenkov light for energy reconstruction and particle identification, we optimized the application of a highly reflective coating. Through viscosity and thickness control, masking, and a custom spray technique, we refined the coating process to reduce surface defects and improve light yield. Together, these efforts strengthen the ICARUS neutrino program and REDTOP's capability of detecting rare decay events.

Visser, Erin [Michigan State U.] (ORCID:0009000184↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

PRISMA: PARALLEL REFINEMENT AND INTEGRATION SYSTEM FOR MULTI-AZIMUTHAL ANALYSIS

The Parallel Refinement and Integration System for Multi-azimuthal Analysis (PRISMA, version 1.1.0) is a Python application for processing X-ray diffraction (XRD) image data. PRISMA wraps GSAS-II to perform azimuthally-binned peak refinement, computes per-frame strain and d-spacing from those fits, and provides three PyQt5 graphical interfaces: (1) a Recipe Builder for selecting GSAS-II control (.imctrl) files, optional mask (.immask) files or threshold-ased masking, reference and experiment image sets, peaks, zimuthal range and bin size, and an optional ceria-based auto-calibration; (2) a Batch Processor that uses Dask on local workstations and pure MPI (mpi4py.futures.MPICommExecutor) on HPC to distribute GSAS-II refinement across cores or compute nodes and write results to a 4-dimensional (peaks x frames x azimuths x measurements) Zarr dataset; and (3) a Data Analyzer that renders heatmaps of fit parameters, strain, frame-to-frame deltas, and percent-change-vs-reference, and exports user-defined subsections to CSV or Excel. The peak-refinement algorithm is deterministic. Benchmark on ALCF Crux: a 20,000-image set, single-peak fit in frame mode with 44 azimuthal bins on 128 nodes x 128 workers, 48 seconds total wall time.

Lorenzo Martin, Maria De La Cinta [Argonne Nationa↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

An MPMD approach coupling electromagnetic continuum mechanics approximations in ALEGRA

In this work, two complementary approximations for describing aspects of continuum electromagnetics in moving media are discussed: electroquasistatic and magnetoquasistatic. Each has been implemented in the finite element shock code ALEGRA for modeling dynamic electromechanical phenomena on typical engineering time scales, with fully integrated circuit coupling. The approximations can be obtained by consistent asymptotic balancing of Maxwell’s equations relative to timescales associated with magnetic diffusion, charge relaxation, and electromagnetic wave propagation. In ALEGRA, the electroquasistatic approximation is used for ferroelectric (FE) modeling, while the magnetoquasistatic approximation is used for magnetohydrodynamic (MHD) modeling. In this paper we introduce for the first time a detailed derivation of a useful quasi-steady “low-R m ” variant of the MHD approximation applicable for cases, such as with detonators, where the thermodynamic pressure arising from Joule heating dominates over magnetic forces. An additional purpose of this paper is to present a coupling mode using Multiple Program-Multiple Data (MPMD) message passing communication that allows the user to run 3D FE problems together with 2D and/or 3D MHD problems with the respective simulation domains coupled through a common circuit equation. The MPMD coupling capability is used here to model the dynamic coupling of a notional ferroelectric generator with an RP-87 exploding bridgewire detonator. The simulated bridgewire heats up and bursts under current generated by simulated depoling of the ferroelectric generator, as a demonstration of the MPMD capability.

42 ENGINEERING↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Synthesis of Correct Digital Controller Models from Specifications by Model Transformation (21-0320)

The design of high consequence controllers (in weapons systems, autonomy, etc.) that do what they are supposed to do is a significant challenge. Testing simply does not come close to meeting the requirements for assurance. Today circuit designers at Sandia (and elsewhere) typically capture the core behavior of their components using state models in tools such as STATEFLOW. They then check that their models meet certain requirements (e.g. “The system bus must not deadlock” or “both traffic lights at an intersection must not be green at the same time”) using tools called model checkers. If the model checker returns “yes” then the property is guaranteed to be satisfied by the model. However, there are several drawbacks to this industry practice: (1) there is a lot of detail to get right, this is particularly challenging when there are multiple components requiring complex coordination (2) any errors returned by the model checker have to be traced back through the design and fixed, necessitating rework, (3) there are severe scalability problems with this approach, particularly when dealing with concurrency. All this places high demands on the designers who now face not only an accelerated schedule but also controllers of increasing complexity. This report describes a new and fundamentally different approach to the construction of safety-critical digital controllers. Instead of directly constructing a complete model and then trying to verify it, the designer can start with an initial abstract (think “sketch”) model plus the requirements, from which a correct concrete model is automatically synthesized. There is no need for post-hoc verification of required functional properties. Having tool to carry this out will significantly impact the nation’s ability to ensure the safety of high-consequence digital systems. The approach has been implemented in a prototype tool, along with a suite of examples, including ones that reflect actual problems faced by designers. Our approach operates on a variant of Statecharts developed at Sandia called Qspecs. Statecharts are a widely used formalism for developing concurrent reactive systems, supporting scalability through allowing state models containing composite states, which are the serial or parallel composition of substates which can themselves contain statecharts. Statecharts enable an incremental style of development, in which states are progressively refined to incorporate greater detail in an incremental model of software development. Our approach formulates a set of constraints from the structure of the models and the requirements and propagates these constraints to a fixpoint. The solution to the constraints is an inductive invariant along with guards on the transitions. We also show how our approach extends to implementation refinement, decomposition, composition, and elaboration. We currently handle safety requirements written in LTL (Linear Temporal Logic)

42 ENGINEERING↗

Hybrid PDES Simulation of HPC Networks Using Zombie Packets

Although high-fidelity network simulations have proven to be reliable and cost-effective tools to peer into architectural questions for high-performance computing (HPC) networks, they incur a high resource cost. The time spent in simulating a single millisecond of network traffic in the highest detail can take hours, even for static, well-behaved traffic patterns such as uniform random. Surrogate models offer a significant reduction in runtime, yet they cannot serve as complete replacements and should only be used when appropriate. Thus, there is a need for hybrid modeling, where high-fidelity simulation and surrogates run side-by-side. Here, we present a surrogate model for HPC networks in which: packets bypass the network, while the network state is left untouched, i.e., suspended. To bypass the network, we use historical data to estimate the arrival time at which every packet should be scheduled at; to suspend the network, all in-flight packets are scheduled to arrive at their destinations, and are kept in the system to awaken as zombies when switching back to high-fidelity. Speedup for a hybrid model is relative to the proportion of surrogate to high-fidelity. This light-weight surrogate obtained up to 76× speedup. Keeping the zombies in the network showed an increase in the accuracy of the high-fidelity simulation on restart when compared to restarting the network from an empty state.

HPC networks↗