Search NASASearch

SEARCH · Search NASA

Results for “Operator inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems demand reliable hazard identification. While large language models (LLMs) have shown promise in automating safety analysis tasks, single-turn, monolithic inference is brittle: it lacks the self-correction, deliberation, and contextual refinement that safety engineers apply iteratively. In this paper, we introduce HAZDIAL, a framework that investigates whether structured agentic dialogue (multi-agent, multi-turn interactions) improves the quality of NLP-based hazard identification over single-pass baselines. We systematically compare two dialogue modalities: adversarial debate and constructive discussion, and propose an genetic algorithm-based agentic interaction optimization. We evaluate all configurations against a curated golden dataset using standard classification metrics (accuracy, precision, recall, F1) and a novel dialogue metrics. This work advances the intersection of dialogue systems, multi-agent reasoning, and AI safety, providing empirical evidence for dialogue-driven hazard analysis.

Das, Sanjay [ORNL] (ORCID:0009000542591915)

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA

Radiochemical transport analysis of gamma spectroscopic data to support estimation of molten salt reactor off-gas inventories

This work introduces a novel application of radiochronometry to estimate nuclide inventories in molten salt reactor off-gas systems based on gamma spectroscopic data from the Molten Salt Reactor Experiment. By analyzing isotopic, isobaric, and isomeric activity ratios, key depletion model parameters related to species transport within the reactor system could be inferred. The findings demonstrate the potential of leveraging a limited subset of gamma spectroscopy measurements to accurately estimate nuclide inventories throughout the off-gas system. The approach can be useful in reactor design activities and support analyses relevant to operations, safety, security, and safeguards.

22 GENERAL STUDIES OF NUCLEAR REACTORS

A Generation-Storage Coordination Dispatch Strategy for Power System Based on Causal Reinforcement Learning

In the backdrop of global energy transformation, power systems integrating high proportions of renewable energy sources are facing unprecedented challenges in operational stability and dispatch efficiency. To address these challenges, this study introduces a generation-storage coordination real-time dispatch strategy based on Causal Power System Dynamic Reinforcement Learning (CPSDRL). Diverging from traditional reinforcement learning approaches, CPSDRL innovatively incorporates causal inference within the state prediction model - the crux of model-based reinforcement learning - thereby establishing the Power Causal Dynamic Model (PCDM). Assisted by the prior knowledge of power systems, the model significantly enhances prediction accuracy and reliability through a two-stage training process. Utilizing PCDM, this study further applies a direct policy search algorithm to optimize the real-time dispatch strategy. Experimental results indicate that the proposed method improves the stability of generation-storage coordination real-time dispatch and exhibits competitive advantages in sample efficiency and computational speed, compared to traditional model-based and model-free reinforcement learning algorithms. This method is expected to enhance the practicality and adaptability of causal reinforcement learning techniques in power system scheduling and control.

causal reinforcement learning

Low-lying baryon resonances from lattice QCD

Recent results studying the masses and widths of low-lying baryon resonances in lattice QCD are presented. The S-wave Nπ scattering lengths for both total isospins I = 1/2 and I = 3/2 are inferred from the finite-volume spectrum below the inelastic threshold together with the I = 3/2 P-wave containing the Δ(1232) resonance. A lattice QCD computation employing a combined basis of three-quark and meson-baryon interpolating operators with definite momentum to determine the coupled channel $Σπ-N\bar{K}$ scattering amplitude in the Λ(1405) region is also presented. Our results support the picture of a two-pole structure suggested by theoretical approaches based on SU(3) chiral symmetry and unitarity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Tokamak divertor plasma emulation with machine learning

Abstract Future tokamak devices that aim to create conditions relevant to power plant operations must consider strategies for mitigating damage to plasma facing components in the divertor. One of the goals of MAST-U tokamak operations is to inform these considerations by researching advanced divertor configurations that aid stable plasma detachment. Machine design, scenario planning and detachment control would all greatly benefit from tools that enable rapid calculation of scenario-relevant quantities given some input parameters. This paper presents a method for generating large, simulated scrape-off layer data sets, which was applied to generate a data set of steady-state Hermes-3 simulations of the MAST-U tokamak. A machine learning model was constructed using a Bayesian approach to hyperparameter optimisation to predict diagnosable output quantities given control-relevant input features. The resulting best-performing model, which is based on a feedforward neural network, achieves high accuracy when predicting electron temperature at the divertor target and carbon impurity radiation front position and runs in around 1 ms in inference mode. Techniques for interpreting the predictions made by the model were applied, and a high-resolution parameter scan of upstream conditions was performed to demonstrate the utility of rapidly generating accurate predictions using the emulator. This work represents a step forward in the design of machine learning-driven emulators of tokamak exhaust simulation codes in operational modes relevant to divertor detachment control and plasma scenario design.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

NLML: A Deep Neural Network Emulator for the Exact Nonlinear Interactions in a Wind Wave Model

Nonlinear wave interactions describe the resonant energy transfer between wave components, playing a fundamental role in the evolution of ocean wave spectra. Nonlinear wave interactions significantly influence wave growth and development, making them essential for accurate wave modeling. However, resolving the full six-dimensional Boltzmann integral of the exact nonlinear wave interactions (Webb-Resio-Tracy method, WRT) is computationally expensive, limiting its application in real-time operational wave forecasting and for research purposes. Current approximations, such as the Discrete Interaction Approximation (DIA), prioritize computational speed over accuracy, resulting in significant errors in wave mean parameters. Here, we introduce NLML, a machine learning (ML) emulator designed to approximate the exact nonlinear wave interactions within WAVEWATCH III (WW3), with the goal of achieving the accuracy of WRT while maintaining the stability and computational speed of DIA. By leveraging GPU capabilities such as half precision inference, we achieved substantial speedups, up to 136x mathematical equation faster than the WRT and only a modest 1.04x mathematical equation slowdown relative to DIA, while achieving 2x mathematical equation the accuracy of DIA in global wave spectral energy and mean wave parameters, with up to 7x mathematical equation higher accuracy in some regions. Unlike previous ML approaches, NLML maintained inherent stability throughout model integration in a standalone, year-long WW3 simulation, without requiring additional constraints. Our new ML parameterization bridges the gap between accuracy and efficiency, offering a promising alternative for improving wave modeling in operational settings and research purposes.

16 TIDAL AND WAVE POWER

Design and simulation of a muon detector to characterize geological overburden

This study presents the design, construction, and simulation of a mobile muon detector tailored for geological overburden characterization. The detector employs plastic scintillator paddles with silicon photomultipliers (SiPMs) and a QuarkNet data acquisition system, offering a portable solution suitable for remote field deployment. The simulator’s modular aluminum frame allows for adjustable geometry and directional sensitivity, while its battery system supports over a week of autonomous operation. Preliminary experimental tests confirmed that its muon flux measurements were consistent with theoretical expectations. A comprehensive simulation framework using Geant4 and CORSIKA was developed to model detector response and overburden effects. Analytical and Monte Carlo methods were used to assess quadrant resolution and infer muon directionality. This work lays the foundation for future overburden mapping and supports the development of reconstruction algorithms for geological applications.

72 - PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

A Unified Interpretation of Variability in Precipitation Isotope Ratios

Abstract Several mechanisms have been proposed to explain why the isotope ratios of precipitation vary in space and time and why they correlate with other climate variables like temperature and precipitation. Here, we argue that this behavior is best understood through the lens of radiative transfer, which treats the depletion of atmospheric vapor transport by precipitation as analogous to the attenuation of light by absorption or scattering. Building on earlier work by Siler et al., we introduce a simple model that uses the equations of radiative transfer to approximate the two-dimensional pattern of the oxygen isotope composition of precipitation ( δ p ) from monthly mean hydrologic variables. The model accurately simulates the spatial and seasonal variability in δ p within a state-of-the-art climate model and permits a simple decomposition of δ p variability into contributions from gradients in evaporation and the length scale of vapor transport. Outside the tropics, δ p is mostly controlled by gradients in evaporation, whose dependence on temperature explains the positive correlation between δ p and temperature (i.e., the temperature effect). At low latitudes, δ p is mostly controlled by gradients in the transport length scale, whose inverse relationship with precipitation explains the negative correlation between δ p and precipitation (i.e., the amount effect). This suggests that the temperature and amount effects are both mostly explained by the variability in upstream rainout, but they reflect distinct mechanisms governing rainout at different latitudes. Significance Statement The isotopic composition of precipitation has long been used to make inferences about past climates based on its observed relationship with precipitation in the tropics and with temperature at higher latitudes. These relationships—known as the “amount effect” and “temperature effect,” respectively—have been attributed to many different mechanisms, most of which are thought to operate at either high or low latitudes but not both. Here, we present a unified framework for interpreting the isotope variability that can explain the latitude dependence of the temperature and amount effects despite making no distinction between high and low latitudes. Although our results are generally consistent with certain interpretations of the amount effect, they suggest that the temperature effect is widely misunderstood.

54 ENVIRONMENTAL SCIENCES

Learning with Adaptive Conservativeness for Distributionally Robust Optimization: Incentive Design for Voltage Regulation: Preprint

Information asymmetry between the Distribution System Operator (DSO) and Distributed Energy Resource Aggregators (DERAs) obstructs designing effective incentives for voltage regulation. To capture this effect, we employ a Stackelberg game-theoretic framework, where the DSO seeks to overcome the information asymmetry and refine its incentive strategies by learning from DERA behavior over multiple iterations. We introduce a model-based online learning algorithm for the DSO, aimed at inferring the relationship between incentives and DERA responses. Given the uncertain nature of these responses, we also propose a distributionally robust incentive design model to control the probability of voltage regulation failure and then reformulate it into a convex problem. This model allows the DSO to periodically revise distribution assumptions on uncertain parameters in the decision model of the DERA. Finally, we present a gradient-based method that permits the DSO to adaptively modify its conservativeness level, measured by the size of a Wasserstein metric-based ambiguity set, according to historical voltage regulation performance. The effectiveness of our proposed method is demonstrated through numerical experiments.

distribution system operator

LUNA: LUT-Based Neural Architecture for Fast and Low-Cost Qubit Readout

Qubit readout is a critical operation in quantum computing systems, which maps the analog response of qubits into discrete classical states. Deep neural networks (DNNs) have recently emerged as a promising solution to improve readout accuracy . Prior hardware implementations of DNN-based readout are resource-intensive and suffer from high inference latency, limiting their practical use in low-latency decoding and quantum error correction (QEC) loops. This paper proposes LUNA, a fast and efficient superconducting qubit readout accelerator that combines low-cost integrator-based preprocessing with Look-Up Table (LUT) based neural networks for classification. The architecture uses simple integrators for dimensionality reduction with minimal hardware overhead, and employs LogicNets (DNNs synthesized into LUT logic) to drastically reduce resource usage while enabling ultra-low-latency inference. We integrate this with a differential evolution based exploration and optimization framework to identify high-quality design points. Our results show up to a 10.95x reduction in area and 30% lower latency with little to no loss in fidelity compared to the state-of-the-art. LUNA enables scalable, low-footprint, and high-speed qubit readout, supporting the development of larger and more reliable quantum computing systems.

Farooq, M. A. [Arizona State U., Tempe]

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja

Power balance and divertor asymmetries in the Super-X divertors of MAST-U using SOLPS-ITER

Spherical tokamaks (STs) present unique challenges and opportunities in the area of particle and power exhaust, intensified due to their more compact sizes. Substantial efforts are underway in STs to determine the limits in dissipative operational regimes and advanced divertor solutions, including at MAST-U which provides access to the Super-X divertor configuration. Power balance, and upper/lower divertor asymmetries have been studied using SOLPS-ITER simulations of the MAST-U Super-X divertor. A set of simulations with experimentally inferred transport coefficients with E x B and diamagnetic drifts activated, consisting of density and power scans, and high field side vs low field side gas puff locations, have been used for code experimentation to uncover trends beyond the current experimental parameter space. The upper biased asymmetry (U:L > 1) of the ratio of the peaks of the plasma energy flux densities at the outer targets increases with heating power and decreases with gas puff strength, going from symmetric to up to a factor of 15. The upper target electron temperature has been found to be a good ordering quantity for the magnitude of this asymmetry for all heating powers, gas puff strength, and gas puff locations. The lower divertor biased asymmetry (U:L < 1) of the radiation patterns processed through SOLPS-based bolometry synthetic diagnostics is in qualitative agreement with resistive bolometry experimental results, and it is in quantitative agreement with the trend of the total volume radiation within the divertors of SOLPS. However, radiation measurements alone are not sufficient to infer the magnitude of the asymmetry of the peaks of the power loads at the targets.

MAST-U

Score-Based Physics-Informed Neural Networks for High-Dimensional Fokker–Planck Equations

The Fokker-Planck (FP) equation is a foundational partial differential equation (PDE) in stochastic processes involving Brownian motions. However, the curse of dimensionality (CoD) poses a formidable challenge when dealing with high-dimensional FP equations. Although Monte Carlo simulation and (vanilla) Physics-Informed Neural Networks (PINNs) have shown the potential to tackle CoD, both methods exhibit significant numerical errors in high dimensions when dealing with the probability density function (PDF) associated with Brownian motion. The point-wise PDF values tend to decrease exponentially as dimensionality increases, surpassing the precision of numerical simulations and resulting in substantial errors. In addition, due to its massive sampling, Monte Carlo fails to offer fast sampling. Modeling the logarithm likelihood (LL) via vanilla PINNs transforms the FP equation into a notoriously difficult Hamilton-Jacobi-Bellman (HJB) equation, which is impractical for PINN learning, whose error grows rapidly with dimension. To this end, we propose a novel approach utilizing a score-based solver to fit the score function in stochastic differential equations (SDEs). The score function, defined as the gradient of the LL, plays a fundamental role in inferring LL and PDF and enables fast SDE sampling, offering an effective means to overcome the CoD. Three fitting methods, Score Matching (SM), Sliced Score Matching (SSM), and Score-PINN, are introduced, each contributing unique advantages in computational complexity, accuracy, and generality. The proposed score-based SDE solver operates in two stages: first, employing score matching or Score-PINN to acquire the score function; and second, solving the LL via an ordinary differential equation (ODE) using the obtained score function. Comparative evaluations across these methods showcase varying trade-offs. The proposed methodology is evaluated across diverse SDEs, including anisotropic Ornstein-Uhlenbeck processes, geometric Brownian motion, and Brownian motion with varying eigenspace. We also test various distributions, including Gaussian, Log-normal, Laplace, and Cauchy distributions. The numerical results demonstrate the score-based SDE solver’s stability, speed, and performance across different experimental settings, solidifying its potential as a solution to CoD for high-dimensional FP equations.

97 MATHEMATICS AND COMPUTING

Acceptance criteria for in situ surveillance of MSR materials based on thermally-loaded mechanical test articles

This report describes practices and acceptance test procedures for designing, running, and maintaining a material surveillance program in a future operating molten salt reactor. The programs described here rely on test data from passively actuated mechanical test articles inserted into critical regions of the reactor and periodically removed for out-of-reactor testing. The report defines definite acceptance procedures, based on the results of these tests, to determine whether a component can continue to operate accounting for the accumulation of environmentally-assisted mechanical damage in the component materials to date, and extrapolated out through the next inspection period. Additionally, the report describes work on a software tool implementing many of the surveillance methods and procedures described here and progress on simplified methods for inferring damage accumulation in the test articles, based on out-of-reactor thermal cycling, that do not rely on sophisticated numerical analysis.

36 MATERIALS SCIENCE

Improving the Confidence in Retrievals of Vertical Distributions of Cloud Condensation Nuclei Number Concentration from ARM Supported by Aircraft In Situ Observations

Accurate quantification of the vertical distribution of cloud condensation nuclei (CCN) number concentrations is critical for improving our understanding of aerosol–cloud interactions. Ground-based Raman lidars operated by the Atmospheric Radiation Measurement (ARM) program, together with surface CCN measurements, are used to retrieve vertically resolved CCN number concentrations (Retrieved Number concentration of CCN, RNCCN). These retrievals rely on several assumptions, including that aerosol composition is vertically homogeneous. To assess this assumption, we developed and tested a framework to infer the dominant aerosol classes/types at different altitudes. This was done by applying a k-Nearest-Neighbors (kNN) algorithm to lidar ratio and linear depolarization ratio measurements from Raman lidar. We evaluated the framework using aircraft aerosol and CCN measurements from the ARM Holistic Interactions of Shallow Clouds, Aerosols, and Land Ecosystems (HI-SCALE) field campaign. The results show that RNCCN performance degrades as vertical aerosol complexity increases, i.e., RNCCN agrees with the aircraft CCN in vertically homogeneous conditions, but closure decreases in layered aerosol structures. To generalize beyond individual examples, we introduce a metric (heterogeneity index) that quantifies the vertical complexity by assessing the variation of inferred aerosol classes/types. Case-level statistics show a tendency for RNCCN and aircraft differences to increase with this metric. By detecting retrievals that are likely compromised by aerosol vertical heterogeneity, the proposed framework improves the interpretability and effective use of RNCCN used for long-term evaluation of models and aerosol–cloud interactions.

Tian, Jingjing