Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Physics-Informed Recurrent Neural Networks to Predict Reactor Operations of the AGN-201 Nuclear Reactor

4 page paper submitted to ANS Student conference. Summary of paper similar to the following abstract: The ability to predict how a reactor will operate, understand when anomalous conditions arise, and ensure a reactor is being operated as expected is crucial for deploying new nuclear facilities. Digital twins serve as a unique solution to recognizing reactor behavior; however, they require data to be useful. For next-generation reactors, this data may not currently be available. To explore how synthetic physics-informed reactor data can be used to predict reactor operations, a recurrent neural network was implemented for the Idaho State University AGN-201 digital twin. The goal of this work is to determine how synthetic data can be used to train a recurrent neural network model for predicting the reactor power of the AGN-201. The recurrent neural network was validated using both synthetic and real operational data. We envision this approach will help bridge the gap between the virtual and physical sides of a digital twin, where reactor physics models based on as-built data can be corrected for actual operating parameters to ensure the virtual model mirrors reality.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

ARPA-E Grid Optimization (GO) Competition Challenge 1

The ARPA-E Grid Optimization (GO) Competition Challenge 1, from 2018 to 2019, focused on the basic Security Constrained AC Optimal Power Flow problem (SCOPF) for a single time period. The Challenge utilized sets of unique datasets generated by the ARPA-E GRID DATA program. Each dataset consisted of a collection of power system network models of different sizes with associated operating scenarios (snapshots in time defining instantaneous power demand, renewable generation, generator and line availability, etc.). The datasets were of two types: Real-Time, which included starting-point information, and Online, which did not. Week-Ahead data is also provided for some cases but was not used in the Competition. Although most datasets were synthetic and generated by GRIDDATA, a few came from industry and were only used in the Final Event. All synthetic Input Data and Team Results for the GO Competition Challenge 1 for the Sandbox, Trial Events 1 to 3, and the Final Event along with problem, format, scoring and rules descriptions are available here. Data for industry scenarios will not be made public. Challenge 1, a minimization problem, required two computational steps. Solver 1 or Code 1 solved the base SCOPF problem under a strict wall clock time limit, as would be the case in industry, and reported the base case operating point as output, which was used to compute the Objective Function value that was used as the scenario score. The feasibility of the solution was provided by the Solver 2 or Code 2, which solves the power flow problem for all contingencies based on the results from Solver 1. This is not normally done in industry, so the time limits were relaxed. In fact, there were no time limits for Trial Event 1. This proved to be a mistake, with some codes running for more than 90 hours, and a time limit of 2 seconds per contingency was imposed for all other events. Entrants were free to use their own Solver 2 or use an open-source version provided by the Competition. Containers, such as Docker, were considered to improve the portability of codes, but none that could reliably support a multi-node parallel computing environment, e.g., MPI, could be found. For more information on the competition and challenge see the "GO Competition Challenge 1 Information" and "GO Competition Challenge 1 Additional Information" resources below.

ACOPF↗

Synthetic Streamflow Datasets to Support Emulation of Water Allocations via LSTM

This archive is the data companion to the bonney_et-al_2026_erc metarepo which generates synthetic data, trains an LSTM model, and generates performance metrics on the trained model. While the generation of the synthetic data is fully reprodicible, it is a computationally expensive process. This data archive contains the synthetic datasets needed for training and testing an LSTM model and reproduction of figures and tables. In addition, supplemenatary data products generating and visualizing results is also included, such as geospatial data for the basin. Contents There are two high level directories: `WRAP_archive/` and `repo_data/`. The `WRAP_archive` directory contains compressed intermediate dataproducts from the dataset generation workflow (marked as "I_Dataset_Generation" in the metarepo). These data products are not required by any scripts in the metarepo, but they are archived as they are expensive to generate and may have useful information for other analyses. The `repo_data` directory contains the necessary data for reproducing the workflow in the metarepo and should be decompressed and moved into the top level of the metarepo. Additional details are provided in README.md.

drought↗

Transfer learning for analysis of collective and non-collective Thomson scattering spectra

Thomson scattering (TS) diagnostics provide reliable, minimally perturbative measurements of fundamental plasma parameters, such as electron density (⁠n e ) and electron temperature (⁠T e ⁠). Deep neural networks can provide accurate estimates of ⁠n e and T e when conventional fitting algorithms may fail, such as when TS spectra are dominated by noise, or when fast analysis is required for real-time operation. Although deep neural networks typically require large training sets, transfer learning can improve model performance on a target task with limited data by leveraging pre-trained models from related source tasks, where select hidden layers are further trained using target data. We present five architecturally diverse deep neural networks, pre-trained on synthetic TS data and adapted for experimentally measured TS data, to evaluate the efficacy of transfer learning in estimating n e and T e in both the collective and non-collective scattering regimes. We evaluate errors in n e and T e estimates as a function of training set size for models trained with and without transfer learning, and we observe decreases in model error from transfer learning when the training set contains ≲ 200 experimentally measured spectra.

Artificial neural networks↗

Anomaly Detection in Seismic Data with Deep Learning: Application for Instrument Failure Detection and Forecasting

Seismic data quality assessment (QA) is the first and one of the most important steps before conducting any further data analysis. Traditional methods involve checking various metrics, such as spike detection and power spectral density, by setting strict thresholds or comparing data against synthetic benchmarks. However, these approaches often rely on pre-existing knowledge and assumptions about data anomalies, leading to potential misclassification of unusual cases. Here, in this study, we propose a deep autoencoder model, an unsupervised learning approach that evaluates data quality without making assumptions about normal and anomalous data, which can be used to identify deviations in recorded data that may indicate nascent instrument failure. We test the model with the U.S. International Monitoring System (IMS) seismic stations and demonstrate the capability of detecting anomalies on a monthly scale. This could prompt station operators to examine potential problems early, allowing sufficient time for instrument maintenance to prevent data outages. In addition, we use a new manually selected testing dataset to compare our model performance against two supervised machine learning (ML) approaches and a standard QA package, as baseline models. When applied to the dataset containing known data anomalies, performance of the supervised and unsupervised ML approaches is similar, with an accuracy of 88.1% for our model compared to ∼90% for the supervised ML approach and 78.2% for the standard QA package. Our model outperforms the baseline models when applied to new stations, where new types of data anomalies can be station-specific and not included in the training dataset. Finally, we show model transferability by training the model with data from the Global Seismograph Network only and applying it to the IMS network data. The results suggest that our model is generalizable and can be applied to new stations with good accuracy.

Lin, Jiun-Ting [Lawrence Livermore National Labora↗

Development of physics-consistent conditional diffusion model to overcome data scarcity in critical heat flux

Deep generative modeling provides a powerful pathway to overcome data scarcity in energy-related applications where experimental data are often limited. By learning the underlying probability distribution of the training dataset, deep generative models, such as the diffusion model, can generate high-fidelity synthetic samples that statistically resemble the training data. Such synthetic data generation can significantly enrich the size and diversity of the available training data, and more importantly, improve the robustness of downstream machine learning models in predictive tasks. The objective of this paper is to investigate the effectiveness of diffusion models for overcoming data scarcity in nuclear energy applications. By leveraging a public dataset on critical heat flux which covers a wide range of commercial nuclear reactor operational conditions, we developed a diffusion model that can generate an arbitrary amount of synthetic samples. Since a vanilla diffusion model can only generate samples randomly, we also developed a conditional diffusion model capable of generating targeted critical heat flux data under user-specified thermal-hydraulic conditions. The performance of the diffusion model was evaluated based on its ability to capture empirical feature distributions and pair-wise correlations, as well as to maintain physical consistency. The results showed that both the diffusion model and conditional diffusion model can successfully generate realistic and physics-consistent critical heat flux data. Furthermore, uncertainty quantification results demonstrate that the conditional diffusion model is highly effective in augmenting critical heat flux data while maintaining acceptable levels of uncertainty.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

Active oversight and quality control in standard Bayesian optimization for autonomous experiments

The fusion of experimental automation and machine learning has catalyzed a new era in materials research, prominently featuring Gaussian Process (GP) Bayesian Optimization (BO) driven autonomous experiments. Here we introduce a Dual-GP approach that enhances traditional GPBO by adding a secondary surrogate model to dynamically constrain the experimental space based on real-time assessments of the raw experimental data. This Dual-GP approach enhances the optimization efficiency of traditional GPBO by isolating more promising space for BO sampling and more valuable experimental data for primary GP training. We also incorporate a flexible, human-in-the-loop intervention method in the Dual-GP workflow to adjust for unanticipated results. We demonstrate the effectiveness of the Dual-GP model with synthetic model data and implement this approach in autonomous pulsed laser deposition experimental data. This Dual-GP approach has broad applicability in diverse GPBO-driven experimental settings, providing a more adaptable and precise framework for refining autonomous experimentation for more efficient optimization.

36 MATERIALS SCIENCE↗

Online and Offline Data Quality Monitoring for the Mu2e Calorimeter

This thesis presents the design, implementation, and validation of a calorimeter Data Quality Monitoring (DQM) toolchain for the Mu2e experiment at Fermilab. Mu2e searches for charged lepton flavor violation via coherent muon-to-electron conversion in the field of an aluminum nucleus, $\mu^- Al \rightarrow e^-Al$, a process whose observation would constitute clear evidence of physics beyond the Standard Model. Achieving target sensitivity requires stringent control of detector performance and data integrity during acquisition, as subtle issues in readout configuration, data formatting, or electronics behavior can compromise reconstruction and bias downstream analyzes. To address these challenges, this work develops a multi-layer DQM approach spanning both raw data validation and reconstructed digi-level diagnostics. At the low level, a fragment analysis component performs word- and bit-field decoding of calorimeter readout blocks, enabling sanity checks of the expected structure and producing detailed error and integrity statistics useful for commissioning and troubleshooting. At the digi level, the CaloDigiDQM analyzer is implemented within the art framework and transforms each CaloDigiCollection into a structured hierarchy of ROOT histograms designed for fast drill-down diagnostics. The module generates coherent monitoring views at global, disk, board, and channel granularity, including occupancy, waveform-derived features (baseline, RMS, peak amplitude and position), and left-right sensor consistency metrics. Detector-aware channel-to-electronics mapping is performed through the conditions system (CaloDAQMap), ensuring that diagnostics remain aligned with hardware identifiers used in operations. For end-to-end testing without reliance on live DAQ data, a synthetic CaloDigi producer is developed to generate realistic waveforms with controlled noise and pulse shapes. The resulting system supports both offline ROOT-file production and online operation, including optional histogram streaming through otsdaq via ots::HistoSender. This toolchain provides a practical and scalable foundation for calorimeter commissioning and stable data collection, enabling early detection of anomalies and reducing operational risk for Mu2e.

Vakulenko, Mark [Drew U.] (ORCID:0009000276197818)↗

Inference of phase field fracture models

The phase field approach to modeling fracture uses a diffuse damage field to represent cracks. This representation mollifies singularities that arise in computations with sharp interface models and some of the resultant difficulties in the mathematical and numerical treatment of fracture. Phase field fracture models have proven effective at representing crack propagation, branching, and merging. Specific formulations, beginning with brittle fracture, have also been shown to converge to classical solutions. Extensions to cover the range of material failure, including ductile and cohesive fracture, lead to an array of possible models. There exists a large body of literature focusing on this class of models and on the impact of model form on the predicted crack evolution. However, there have not been systematic studies into how optimal models may be chosen. Here, we take a first step in this direction by developing formal methods for identification of the best parsimonious model of phase field fracture given full-field data on the damage and deformation fields. We consider some of the main models that have been used for the degradation of elastic response due to damage and its propagation. Our approach builds upon Variational System Identification (VSI), a weak form variant of the Sparse Identification of Nonlinear Dynamics (SINDy). Furthermore, in this first communication we focus on synthetically generated data but we also consider central issues associated with the use of experimental full-field data, such as data sparsity and noise.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Learning Constitutive Relations From Soil Moisture Data via Physically Constrained Neural Networks

Abstract The constitutive relations of the Richardson‐Richards equation encode the macroscopic properties of soil water retention and conductivity. These soil hydraulic functions are commonly represented by models with a handful of parameters. The limited degrees of freedom of such soil hydraulic models constrain our ability to extract soil hydraulic properties from soil moisture data via inverse modeling. We present a new free‐form approach to learning the constitutive relations using physically constrained neural networks. We implemented the inverse modeling framework in a differentiable modeling framework, JAX, to ensure scalability and extensibility. For efficient gradient computations, we implemented implicit differentiation through a nonlinear solver for the Richardson‐Richards equation. We tested the framework against synthetic noisy data and demonstrated its robustness against varying magnitudes of noise and degrees of freedom of the neural networks. We applied the framework to soil moisture data from an upward infiltration experiment and demonstrated that the neural network‐based approach was better fitted to the experimental data than a parametric model and that the framework can learn the constitutive relations.

54 ENVIRONMENTAL SCIENCES↗

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING↗

ARPA-E Grid Optimization (GO) Competition Challenge 3

Synthetic Input Data and Team Results for the GO Competition Challenge 3 for Events 1 - 4 and the Sandbox, along with problem and format descriptions and code to validate data and solutions, are available here. Data for industry scenarios will not be made public. The Grid Optimization (GO) Competition Challenge 3 focused on the security-constrained optimal power flow (SCOPF) problem. It is part of a continuing effort begun with Challenges 1 and 2, to successfully discover, develop, and test innovative and disruptive software solutions for critical energy challenges and to overcome existing barriers. The broader goal of the of the GO Competition is to accelerate the development of transformational and disruptive methods for solving problems related to the electric power grid and to provide a transparent, fair, and comprehensive evaluation of new solution methods. Challenge 3 used multiperiod dynamic markets, including advisory models for extreme weather events, day-ahead markets, and the real-time markets with an extended look-ahead. In Event 4, whose submission window was August 31-September 4, 2023, 14 teams solved for the objective values of 669 scenarios (39 scenarios required solutions both with and without line switching being allowed). The 591 synthetic scenarios from 9 network models (3.6 GB) are available here. Ten teams were funded to participate and 7 won prizes totaling $2,400,000. The largest prize ($550,000) went to Mississippi State University. An additional $600,000 was awarded in Event 3 (6/15-16/2023). No prizes were awarded in Events 1 (1/25-27/2023) or 2 (4/13-14/2023). For more information on the competition and challenge see the "GO Competition Challenge 3 Information" resource below.

ACOPF↗

Simulation of Physics-Based 0-10Hz Strong Motion Using High Performance Computing Supporting Refinements to Regional Ground Motion Models for the Central Eastern US

In collaboration with the U.S. Nuclear Regulatory Commission (NRC) the LLNL has developed a computationally efficient simulation platform designed to perform physics-based ground motion simulations for crustal earthquakes in the Stable Continental Regions of Central and Eastern US (CEUS), using high-performance computing. The main objective of the earthquake simulations was to use synthetic ground motion to provide constrains to refinements of existing ergodic Ground Motion Models (GMMs), for large magnitude earthquakes and near-fault distances, for which these models are less reliable. Physics-based broadband (0-10Hz) ground motion simulations were used to estimate the near-fault ground motion amplitudes and within event and between-event variabilities associated with fault rupture characteristics. In our simulations we used a 3D regional velocity model that was based on Saikia’s 1D velocity model (1994). In simulations performed during the first stage of this project the Saikia’s velocity model demonstrated better performance in modelling high frequency regional wave propagation for the CEUS region recorded during the Mw5.0 November 7, 2016, Cushing Oklahoma (Taylor et al., 2017), and Mw5.8 September 3, 2016, Pawnee Oklahoma earthquakes. The proposed regional 3D model includes random perturbations to the 1D background model using the stochastic scheme of Pitarka and Mellors (2021). In addition, validation analysis of the rupture generator and regional wave propagation models, using comparisons with different GMMs for Mw6.5 and Mw7.0 scenario earthquakes in the CEUS region resulted in a very good match between the simulated and empirical ground motion models. For the purposes of seismic hazard assessment at the existing and planned nuclear power plants, NRC is interested in studies aimed at improving the current ground motion models (GMM) for both Stable Continental Regions (SCR) in the Central and Eastern US and Active Crustal Regions (ACR) in the Western US. Due to lack of recorded data, these improvements require synthetic data for short fault distances and large magnitude earthquakes for which the existing recorded data is not enough to uniquely constrain the GMMs. The need for simulations and strong motion data is especially critical for the CEUS region where we do not have recorded data from potentially large damaging earthquakes with moment magnitudes 6.0 and higher. In this the project, we focused on 10Hz simulations of Mw7.0 scenario earthquakes with strike slip and thrust faulting mechanisms. We used more than 50 Mw7.0 earthquake rupture scenarios to investigate the ground motion uncertainty due to unknown earthquake rupture parameters, in particular, the slip distribution, rupture velocity, and faulting mechanism, and their implication on ground motion amplification due to forward rupture directivity effects.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Velocity Extraction Using Complete Time-Domain Waveform Data and Audio Machine Learning

We developed a new machine learning-based tool for extracting information from interferometry measurements: MIDWAZE (Modular Interferometry Direct Waveform AnalyZEr). This paper showcases MIDWAZE’s ability to extract an object’s velocity information from Photonic Doppler Velocimetry (PDV) data at near-human accuracy with little to no human intervention. MIDWAZE can extract velocities roughly 350 times as fast as a human analyst "rushing" to complete their extractions, with similar extraction accuracy. MIDWAZE’s most outstanding feature is that it operates directly in waveform/temporal space, freeing analysis from certain limitations imposed by traditional spectrogram-based approaches and opening the way to "phase aware" PDV analysis. MIDWAZE also has limited ability to discriminate between different solid objects, which we develop as a first step towards automated discrimination of different kinds of objects such as ejecta clouds.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multi-objective optimization of sustainable aviation fuel production pathways in the U.S. Corn Belt

As a potential source of low-carbon transportation energy, biofuels offer certain advantages over vehicle electrification (e.g., lower societal vulnerability to grid failures, and improved range of sustainable aviation), but also several challenges, including cost, carbon intensity, and land usage. There are also well-founded concerns that biofuel supply chains could be disrupted if extreme weather events impact feedstock yields. In this paper, we explore the use of multi-objective optimization to identify biofuel production pathways that balance cost, greenhouse gas emissions, and supply vulnerability to extreme weather. We compare the use of three different many-objective evolutionary algorithms and linear programming in optimizing biomass cultivation decisions in the U.S. Corn Belt under weather uncertainty using historical, modeled, and synthetic yield data. We consider four feedstock choices (corn, soy, switchgrass, and algae) with two land types (agricultural and marginal lands) and evaluate decisions using three alternative spatial resolutions (ranging from the USDA agricultural district level to the state level). Results show that feedstock choice is the primary driver of objective performance (i.e., the position and shape of 3D, approximate Pareto frontiers). Spatial diversification is a less effective tool in reducing exposure to weather-caused drops in crop yield.

09 BIOMASS FUELS↗

The use of digital thread for reconstruction of local fiber orientation in a compression molded pin bracket via deep learning

A deep convolutional neural network (DCNN) was used for microstructure reconstruction using artificial intelligence (MR-AI) by predicting local average fiber orientation distributions (FOD) in a 3D prepreg platelet molded composite (PPMC) pin bracket. To train the MR-AI model, surface strain fields from residual stresses simulated in PPMC plates were used as the input to the DCNN. A training dataset included PPMC plates with various degrees of global fiber alignment, based on the information obtained from high-fidelity flow simulation of a pin bracket. Further, the MR-AI model was then deployed to analyze FOD in the 3D pin bracket by conducting thermo-elastic residual stress analysis. Initially, the MR-AI model was established entirely on the synthetic simulation data. Then, a μCT scan of a physically molded pin bracket was used to create a finite element model that provided data for additional validation of the DCNN model. For the μCT scan finite element pin bracket the MR-AI model predicted the distribution of fiber orientation tensor components with MAE of 0.10 indicating a global prediction error of 10%. For the flow simulated pin bracket, the MR-AI model predicted the distribution of fiber orientation tensor components with a global prediction error of 11%. The MR-AI model showed the ability to predict regions of varying alignment in the base and flange of the pin bracket. The proposed MR-AI methodology allows for rapid prediction of FOD in geometrically complex parts and offers a promising path to detecting unique fiber orientation states in molded components.

42 ENGINEERING↗