Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Development of a Discrepancy Checker for the Digital Twin in a Supervisory Control System for a Thermal Energy Delivery System

Defined as a virtual representation of a physical object, process, or service, and used to support real-world decision-making, a digital twin (DT) can be utilized to combine classical and novel frameworks in sensors, state predictions, and multi-input/multi-output systems, and to enable optimal autonomous operations. However, a DT’s usefulness largely depends on its ability to adequately mirror the state of its physical counterpart, and this adequacy should be reflected by the level of uncertainty in the underlying simulation models when estimating and predicting quantities of interest (QOIs). Moreover, simulation models in a DT may involve multiple fidelities of representations—ranging from physics-based models to data-driven ones—but classical uncertainty quantification (UQ) methods struggle to handle numerous uncertainty sources, nor are they designed for real-time applications. This work presents a UQ-based discrepancy checking and diagnosis tool for a DT-based supervisory control system applied to a thermal energy delivery system (TEDS) at Idaho National Laboratory. The discrepancy checker was developed using metadata from an automated DT development process, and these metadata included different combinations of physical model forms and model parameters, training data and hyperparameters for surrogate models, and design parameters for supervisory control systems. Next, correlations between the uncertainty results and the metadata were established and then applied to the DT operations. The discrepancy checker evaluates the discrepancies between model predictions from virtual and sensor measurements and backtraces them to the corresponding major sources of uncertainty. The discrepancy checker showed reasonable performance in detecting discrepancies and diagnosing sources of uncertainty in testing scenarios.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Predictive Modeling and Uncertainty Quantification in Condition Monitoring of Active Components: A Reactor Coolant Pump Use Case

This work develops data-driven models for onset of thermal barrier leakage in reactor coolant pumps. It incorporates uncertainty quantification to enhance the reliability and robustness of pre- dictions. Using synthetic data generated by the Generic Pressurized Water Reactor simulator, realistic degradation scenarios were simulated across lifecycle stages—beginning, middle, and end of life. Key variables, including differential pressure, flow rate, vibration, and temperatures, were analyzed using machine learning framework. The fully connected neural network models demonstrated exceptional performance, achieving R2 scores exceeding 0.99 and root mean square errors as low as around 8.23 × 10-2 gallon per minute (gpm) for the three stages of the lifecy- cle. UQ analysis further validated the model’s robustness, with narrow uncertainty bounds during steady-state operations and appropriately wider bounds during transitional phases, reflecting the physical behavior of the system. This work addresses important gaps in real-time condition moni- toring and regulatory compliance by integrating advanced condition monitoring technologies with UQ into IST programs. The ability to detect thermal barrier leakage early and quantify prediction reliability supports optimizing maintenance strategies while ensuring nuclear power plants’ safe and reliable operation.

99 - GENERAL AND MISCELLANEOUS↗

A Study on Co-existing Heterogeneous Wireless Networks for Data Transmission within a Nuclear Facility

Deployment of wireless technologies is a salient need for modernization, automation and improved operation of nuclear power plants (NPPs). As a single technology cannot support the ever-changing needs, it is required to have a heterogeneous wireless network architecture to address the different technical and economic challenges. However, the coexistence of these multiband heterogeneous wireless networks brings numerous challenges due to the factors including dissimilarity in their channel access mechanism, distance between nodes, transmit power level and many more. This paper develops real-world experiments and simulations of wireless coexistence for Wi-Fi, Fifth generation cellular (5G) and Zigbee in the unlicensed band to understand the challenges and opportunities. The experiments were conducted over the Platform for Open Wireless Data-driven Experimental Research (POWDER) testbed at the university of Utah. In addition, this paper is the first to propose a novel packet rate control technique at the network layer to create temporary opportunities for 5G or Zigbee signal transmissions focusing its application in a nuclear facility while using the shared band. The performance of the proposed coexistence solution is validated with experimental results and simulation.

5G↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗

Flavor as an Incomplete Structure: Conceptual Questions and the Role of DUNE

Flavor remains one of the most successful yet least understood structures of the Standard Model. The discovery of the Higgs boson completed the electroweak account of mass generation, but did not explain the origin of fermion families, mass hierarchies, or mixing patterns. In this sense, flavor can be regarded as an empirically successful but conceptually incomplete structure. Neutrinos occupy a particularly sensitive place within this problem: their masses are tiny, their mixing is large, and their mass-generation mechanism may differ from that of charged fermions. In this article, we discuss flavor as an open conceptual problem and argue that DUNE, as a phased program spanning precision oscillation measurements and sensitivity to BSM and dark-sector phenomena, provides a powerful framework for testing the self-consistency and possible limits of the present three-flavor description. In particular, the complementarity between the long-baseline program and the Phase I near-detector complex, together with the DUNE-PRISM strategy for controlling interaction-model systematics and enabling data-driven near-to-far predictions, makes DUNE especially well-suited to search for small, correlated departures from the minimal flavor framework.

Montanari, Claudio S. [Fermilab; INFN, Pavia] (ORC↗

Measuring the neutrino-oxygen neutral current quasielastic cross section using the accelerator neutrino neutron interaction experiment

The Accelerator Neutrino Neutron Interaction Experiment (ANNIE) is a 26-ton gadolinium-doped water Cherenkov detector located on-axis to Fermilab’s Booster Neutrino Beam (BNB). ANNIE is uniquely positioned to perform high-statistics measurements of neutrino-nucleus interactions in water, benefiting from a large neutrino flux due to a short (100-meter) baseline. A central focus of ANNIE’s physics program is the measurement of both charged current (CC) and neutral current (NC) cross sections on water, including neutral current quasielastic (NCQE) and CC-inclusive channels. The NCQE measurement is particularly critical for constraining uncertainties in rare-event searches such as the Diffuse Supernova Neutrino Background (DSNB), where atmospheric $\nu$NCQE interactions constitute a significant and poorly constrained background. This dissertation presents a measurement of the flux-averaged neutrino-oxygen neutral current quasielastic ($\nu$NCQE) cross section using $2.573 \times 10^{20}$~POT of BNB exposure from the 2022 and 2023 beam years. The $\nu$NCQE interaction is identified through the primary $\gamma$-rays produced by nuclear de-excitation of the residual $^{15}$N$^*$ or $^{15}$O$^*$ nucleus following nucleon knockout from $^{16}$O. A dedicated Monte Carlo (MC) re-tuning campaign was conducted using an americium-beryllium (AmBe) calibration source, Michel electrons from stopped muons, and throughgoing dirt muons originating upstream of the detector. This multi-sample approach provided a wide-ranging $\mathcal{O}(\text{MeV})$--$\mathcal{O}(\text{GeV})$ dataset for tuning the simulated detector response, which was subsequently validated against AmBe neutron and Michel electron data for use in the $\nu$NCQE analysis. A dedicated laser calibration campaign was carried out to reduce timing uncertainties across the PMT system, enabling reconstruction of the BNB bunch substructure with sufficient resolution to serve as a background rejection tool. By selecting events in-time with individual neutrino bunches, beam-correlated $\nu$NCQE events are separated from diffuse and accelerator-induced backgrounds, notably skyshine neutrons and externally-originating events, that would otherwise dominate traditional charge-based selections within a small-scale, surface-level, short-baseline detector. A data-driven estimation of the skyshine neutron and external background rates was performed and incorporated into the systematic uncertainty budget. The flux-averaged $\nu$NCQE cross section on oxygen is measured to be $1.57 \pm 0.06\,(\text{stat.})$ $^{+0.91}_{-0.67}\,(\text{syst.})$ $\times 10^{-38}\ \text{cm}^{2}$. A full systematic budget is constructed by propagating uncertainties in the secondary hadronic interaction modeling, background cross section normalizations, detector response, neutrino flux, and the primary $\gamma$-ray emission probabilities from oxygen nuclear de-excitation. An idealized de-excitation model, constructed from existing measurements in the literature is developed to benchmark the predictions of the \textsc{GENIE} event generator. A comparison reveals that \textsc{GENIE} systematically overpredicts the primary $\gamma$-ray emission probability from oxygen de-excitation by a factor of $1.49\times$ for $E_\gamma > 6$~MeV and $3.07\times$ in the $3$--$6$~MeV band. This comparison motivates the dominant systematic uncertainty in this analysis, where a conservative uncertainty of $^{+39.9\%}_{-0\%}$ on the primary $\gamma$-ray signal prediction is assigned. The ANNIE result is consistent with and complementary to existing flux-averaged $\nu$NCQE cross section measurements from T2K and Super-Kamiokande, providing an independent measurement with a different detector, neutrino beam, and analysis methodology. Looking ahead, an upgrade to the ANNIE DAQ infrastructure enabling continuous extended readout will allow a complementary $\nu$NCQE neutron multiplicity measurement, directly relevant to constraining the NCQE background in DSNB searches, competitive with the recent T2K measurement at SK-Gd. The planned Super-SANDI upgrade, deploying a large Water-based Liquid Scintillator (WbLS) volume, will further extend ANNIE's reach to hadronic final states and exclusive NC channels, and enable joint measurements with liquid argon detectors sharing the BNB beamline ahead of DUNE and Hyper-Kamiokande.

Doran, Steven [Iowa State U.]↗

Standard Candles for Supernova Neutrino Detection at DUNE

The Deep Underground Neutrino Experiment (DUNE) far detector is sensitive to $\mathcal{O}(10)$ MeV electron neutrinos through $ν_e$ charged-current reaction with argon. This capability is a unique window into the $ν_e$ component of a Galactic core-collapse supernova flux. Extracting the properties of the supernova spectrum is, however, limited by the poorly-known $ν_e$-Ar cross section. We propose a data-driven strategy that leverages $^8$B solar neutrinos and muon-decay-at-rest neutrinos as standard candles for this process. These calibration samples constrain both the low-energy and high-energy components of the cross section relevant for supernova detection. Our method reduces the reliance on nuclear models, which can bias the extraction of the spectral parameters by as much as 300$\%$.

Cheng, Ting [Fermilab]↗

Evaluation of a Reduced-Order Model for IBR Fault Response Representation via OEM Blackbox Models: Preprint

Driven by the need to capture the electromagnetic transients of transmission lines, inverter switching behavior, and detailed control systems, electromagnetic transient (EMT) studies have become increasingly important in industry, such as IBR interconnection study and fault study. However, original equipment manufacturer (OEM) inverter models typically include extensive parameters and proprietary settings that are unavailable to protection engineers. This paper introduces a data-driven, reduced-order model (ROM) developed as a PSCAD library component for use in EMT-based fault studies. The ROM replicates key OEM model behaviors without requiring detailed knowledge of control design or parameterization. The accompanying Python automation scripts streamline data generation, parameter fitting, and validation. The ROM's performance is demonstrated through comparison with both IEEE 2800-compliant and non-compliant OEM models in a real-world power system. Relay responses show nearly identical results, while simulation runtime is reduced by an average of 32.8\%, highlighting the ROM's practicality for protection engineers.

14 SOLAR ENERGY↗

Virtual Growth of SRF Materials: A Machine Learning Approach to Predict the Crystalline Structural Ordering in Nb Surface Oxides

Niobium's native surface oxide affects SRF cavity and superconducting qubit performance, motivating interest in controlling its crystalline structure. We combine a literature-derived machine-learning analysis with temperature-dependent XRD to study crystalline ordering in Nb2O5. Random Forest models, trained on 74 processing conditions from 17 papers and validated by leave-one-group-out cross-validation, predicted broad crystallinity outcomes well (balanced accuracy 0.809), but struggled with specific polymorph identity (0.577). Annealing temperature was the dominant predictor across all targets; oxygen partial pressure showed negligible importance, reflecting narrow literature coverage rather than physical irrelevance. Temperature-dependent XRD on anodized and H2O2-treated Niobium showed structural evolution consistent with the machine learning predictions. Our model and overall approach provide a data-driven framework for identifying and optimizing conditions that promote crystallization in initially amorphous oxides. This framework can guide the selection of growth and post-annealing conditions for Nb surfaces by narrowing the experimental parameter space, thereby reducing trial-and-error efforts in developing oxide structures relevant to SRF applications.

Tilkin, Anthony [Unlisted, US, IL; Fermilab]↗

Determining the Efficiency of EMPHATICs Silicon Strip Detectors (SSDs)

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, V. [Illinois U., Urbana (main)]↗

Neutron detector response modeling in NOvA

Neutrons can present a significant challenge for neutrino experiments in which energy reconstruction is critical. With the ability to escape detection completely and with a weak correlation between their kinetic energy and any eventual energy deposition, it is difficult to fully account for neutrons produced in neutrino interactions. This in turn leads to significant model dependence when evaluating neutron-related systematic uncertainties. The NOvA experiment is a long-baseline neutrino oscillation experiment with a high-statistics sample of antineutrino data collected by its near detector. We report an excess relative to data of simulated neutron candidates with low energy depositions when using standard Geant4 physics lists. The simulation excess is traced to an overabundance of secondary photons produced from interactions of neutrons with kinetic energy greater than \SI{20}{\mega\eV}. Improved agreement with data is obtained by applying the data-driven neutron-on-carbon \menate model for neutrons between \SI{20}{\mega\eV} and ${\sim}$\SI{100}{\mega\eV}. With \menate, the residual oversimulation is more uniform across the calorimetric neutron energy spectrum, suggesting possible overproduction of primary neutrons by the GENIE neutrino interaction generator. These results motivate the adoption of \menate-supplemented Geant4 simulation as the nominal simulation in the production of future \nova simulation.

Abubakar, S.↗

Pre-training Vision Models for the Classification of Alerts from Wide-field Time-domain Surveys

Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.

79 ASTRONOMY AND ASTROPHYSICS↗

Identifying Differential Equations in Fourier Domain (FourierIdent)

We investigate identifying differential equations in the frequency domain. Fourier analysis is an important tool in theoretical analysis and numerical solvers of differential equations, yet there is limited work in exploring this connection in the identification of differential equations. This paper aims to identify the underlying differential equation in the frequency domain, from a given single realization of the differential equation perturbed by noise. Such setting imposes difficulties which are different from other identification methods where computation is carried out in the physical domain. We propose several ways to mitigate the challenges arising from noise in data and large differences in the magnitudes of frequency responses. The main takeaways are that identifying differential equations solely in the frequency domain is challenging, the method we propose is based on a form of domain partitions in the frequency domain, and this method shows benefits for complex data even with high level of noise. We introduce a Fourier feature denoising, and define the meaningful data region and the core regions of features to reduce the effect of noise in the frequency domain and to enhance the accuracy in coefficient identification. The proposed method is tested on various differential equations with linear, nonlinear, and high-order derivative feature terms, and shows advantages on complex data with many frequency modes, even under high level of noise.

97 MATHEMATICS AND COMPUTING↗

Chapter 4 - Recent Advances in Identification of Differential Equations from Noisy Data: IDENT Review

Differential equations and numerical methods are extensively used to model various real-world phenomena in science and engineering. With modern developments, we aim to find the underlying differential equation from a single observation of time-dependent data. If we assume that the differential equation is a linear combination of various linear and nonlinear differential terms, then the identification problem can be formulated as solving a linear system. The goal then reduces to finding the optimal coefficient vector that best represents the time derivative of the given data. We review some recent works on the identification of differential equations. We find some common themes for the improved accuracy: (i) The formulation of linear system with proper denoising is important, (ii) how to utilize sparsity and model selection to find the correct coefficient support needs careful attention, and (iii) there are ways to improve the coefficient recovery. We present an overview and analysis of recent developments on the topic.

97 MATHEMATICS AND COMPUTING↗

A filter-dependent granular temperature model from large-scale CFD-DEM data

The computational study of strongly-coupled, gas–solid flows at scales relevant to most environmental and engineering applications requires the use of ‘coarse-grained’ methodologies such as the two-fluid model, particle-in-cell approach or the multiphase Reynolds Averaged Navier–Stokes equations. While these strategies enable computations at desirable length- and time-scales, they rely heavily on models to capture important flow physics that occur at scales smaller than the mesh. To date, the models that do exist are based on a limited set of flow conditions, such as very dilute particle phase. To this end, we leverage a large-scale repository of CFD-DEM data to develop filter-size dependent models for the mean variance in particle volume fraction, a quantity commonly used to assess the degree of clustering, and the granular temperature, a key quantity for accurately predicting gas–solid flows. In conclusion, because of its filter-size dependence, the granular temperature model can be directly translated to coarse-grained approaches and tied directly to grid size.

AMReX↗

Regional surrogates for predictive control of digital twins

Digital twins of complex systems must involve a model that is fast, generalizable, and usable for real-time control. For example, high-fidelity nonlinear multiphysics simulations can capture laser-material interactions, but are too slow for optimization or model predictive control (MPC). Reduced-order models, used to accelerate such computation, frequently fail to generalize to unseen inputs or control states. We show theoretically that this failure is intrinsic, i.e., that a learned model is non-unique outside the sampled subspace when its low-rank structure arises from limited excitation and clustered eigenvalues, rather than from a user-imposed truncation alone. Motivated by this result, we propose a control-ready regional surrogate-construction framework for both autonomous and nonautonomous dynamics; it employs Koopman lifting to represent nonlinearities, while preserving spatial locality. We illustrate our approach by constructing a control-ready surrogate for the digital twin of a thermal component of additive-manufacturing process. Our surrogate, localized in space through a von Neumann stencil, is learned from noisy high-fidelity simulations that emulate thermal-camera images collected during the manufacturing. It is linear in thermo-physically augmented states so that MPC reduces to a convex quadratic program. The surrogate requires no online correction, generalizes to unseen scan paths and power profiles of the laser, and is more than three orders of magnitude faster than a finite-difference solver. Furthermore, when the MPC sequence computed on the digital twin is applied to this solver, closed-loop temperature regulation is recovered, showing that the surrogate preserves control-relevant input-output behavior.

Data-driven model↗

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗