Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Kimberlina 1.2 CCUS Geophysical Models and Synthetic Data Sets

This synthetic multi-scale and multi-physics data set was produced in collaboration with teams at the Lawrence Berkeley National Laboratory, National Energy Technology Laboratory, Los Alamos National Laboratory, and Colorado School of Mines through the Science-informed Machine Learning for Accelerating Real-Time Decisions in Subsurface Applications (SMART) Initiative. Data are associated with the following publication: Alumbaugh, D., Gasperikova, E., Crandall, D., Commer, M., Feng, S., Harbert, W., Li, Y., Lin, Y., and Samarasinghe, S., “The Kimberlina Synthetic Geophysical Model and Data Set for CO2 Monitoring Investigations”, The Geoscience Data Journal, 2023, DOI: 10.1002/gdj3.191. The dataset uses the Kimberlina 1.2 CO2 reservoir flow model simulations based on a hypothetical CO2 storage site in California (Birkholzer et al., 2011; Wainwright et al., 2013). Geophysical properties models (P- and S-wave seismic velocities, saturated density, and electrical resistivity) were produced with an approach similar to that of Yang et al. (2019) and Gasperikova et al. (2022) for 100 Kimberlina 1.2 reservoir models. Links to individual resources are provided below: [CO2 Saturation Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-co2-saturation-models); Resistivity Models – [part 1](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-1), [part 2](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-2), and [part 3](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-3); [Vp Velocity Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-vp-velocity-models); [Vs Velocity Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-vs-velocity-models); [Density Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-density-models). The 3D distributions of geophysical properties for the 33 time stamps of the SIM001 model were used to generate synthetic seismic, gravity, and electromagnetic (EM) responses for 33 times between zero and 200 years. Synthetic surface seismic data were generated using 2D and 3D finite-difference codes that simulate the acoustic wave equation (Moczo et al., 2007). 2D data were simulated for six point-pressure sources along a 2D line with 10 m receiver spacing and a time spacing of 0.0005 s. 3D simulations were completed for 25 surface pressure sources using a source separation of 1 km in both the x and y directions and a time spacing of 0.001 s. Links to individual resources are provided below: [2D velocity models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-2d-velocity-models) and [2D surface seismic data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-2d-surface-seismic-data). [3D velocity models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-velocity-models), and 3D seismic data [year0](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year0), [year1](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year1), [year2](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year2), [year5](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year5), [year10](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year10), [year15](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year15), [year20](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year20), [year25](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year25), [year30](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year30), [year35](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year35), [year40](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year40), [year45](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year45), [year49](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year49), [year50](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year50), [year51](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year51), [year52](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year52), [year55](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year55), [year60](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year60), [year65](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year65), [year70](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year70), [year75](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year75), [year80](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year80), [year85](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year85), [year90](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year90), [year95](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year95), [year100](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year100), [year110](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year110), [year120](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year120), [year130](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year130), [year140](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year140), [year150](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year150), [year175](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year175), [year200](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year200). The Python scripts to read these models and data are provided [here](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-python-scripts). EM simulations used a borehole-to-surface survey configuration, with the source located near the reservoir level and receivers on the surface using the code developed by Commer and Newman (2008). Pseudo-2D data for the source at [2500 m](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-pseudo-2d-csem-data-tz2500m) and [3025 m](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-pseudo-2d-csem-data-tz3025m), used a 2D inline receiver configuration to simulate a response over 3D resistivity models. The [3D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-csem-data) contain electric fields generated by borehole sources at monitoring well locations and measured over a surface receiver grid. Vector gravity data, both on the surface and in boreholes, were simulated using a modeling code developed by Rim and Li (2015). The simulation scenarios were parallel to those used for the EM: [pseudo-2D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-gravity-data) were calculated along the same lines and within the same boreholes, and [3D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-gravity-data) were simulated over 3D models on the surface and in three monitoring wells. A series of [synthetic well logs](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-well-logs) of CO2 saturation, acoustic velocity, density, and induction resistivity in the injection well and three monitoring wells are also provided at 0, 1, 2, 5, 10, 15, and 20 years after the initiation of injection. These were constructed by combining the low-frequency trend of the geophysical models with the high-frequency variations of actual well logs collected in the Kimberlina 1 well that was drilled at the proposed site. Measurements of permeability and pore connectivity were made on cores of Vedder Sandstone, which forms the primary reservoir unit: [CT micro scans](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-ct-micro-scans-of-vedder-formation) and [Industrial CT Images](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-industrial-ct-images-vedder-formation). These measurements provide the range of scales in the otherwise synthetic data set to be as close to a real-world situation as possible. References: Birkholzer, J.T., Zhou, Q., Cortis, A. and Finsterle, S., 2011. A sensitivity study on regional pressure buildup from large-scale CO2 storage projects. Energy Procedia, 4, 4371-4378. Commer, M., and Newman, G.A., 2008. New advances in three-dimensional controlled-source electromagnetic inversion, Geophysical Journal International, 172, 513-535. Gasperikova, E., Appriou, D., Bonneville, A., Feng, Z., Huang, L., Gao, K., Yang, X., Daley, T., 2022, Sensitivity of geophysical techniques for monitoring secondary CO2 storage plumes, Int. J. Greenh. Gas Control, Volume 114, 103585, ISSN 1750-5836, https://doi.org/10.1016/j.ijggc.2022.103585. Moczo, P., J.O. Robertsson and L. Eisner, 2007, The finite-difference time-domain method for modeling of seismic wave propagation: Advances in geophysics, 48, 421-516. Rim, H., and Y. Li, 2015, Advantages of borehole vector gravity in density imaging, Geophysics, 80, G1-G13. Wainwright, H. M.; Finsterle, S.; Zhou, Q.; Birkholzer, J. T., 2013. Modeling the Performance of Large-Scale CO2 Storage Systems: A Comparison of Different Sensitivity Analysis Methods. International Journal of Greenhouse Gas Control, 17, 189205. https://doi.org/10.1016/j.ijggc.2013.05.007, DOI: 10.18141/1603331. Yang, X., Buscheck, T.A., Mansoor, K., Wang, Z., Gao, K., Huang, L., Appriou, D., and Carroll, S.A., 2019. Assessment of geophysical monitoring methods for detection of brine and CO2 leakage in drinking water aquifers, International Journal of Greenhouse Gas Control, 90, 102803, https://doi.org/10.1016/j.ijggc.2019.102803.

CCUS↗

Synthetic Data Generation Using Machine Learning

Robust machine learning techniques for image analysis require a substantial amount of data to yield confident results. In the nuclear domain, data scarcity is a substantial challenge because there are so few facilities worldwide. This research focuses on being able alleviate the data scarcity problem by generating synthetic data to bridge the gap between large and small datasets. This work achieves that goal using a Generative Adversarial Network (GAN) architectural approach, by training a model on real-world data and expands that small dataset through synthetic data amendments. Model performance is impacted by the size of the real-world dataset and the number of training epochs utilized. This means that 1) It is important to develop your GAN to be optimized with the specific data type, and 2) approaches taken when training the GAN should be specialized to encompass important aspects of the dataset that it is generating. By taking a step to improve dataset sizes in this way, the gap between models trained by parties with significant amount of data and those without access to large data, closes, allowing for robust analyses of satellite imagery for nuclear domain applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Validating automated resonance evaluation with synthetic data

The integrity and precision of nuclear data are crucial for a broad spectrum of applications, from national security and nuclear reactor design to medical diagnostics, where the associated uncertainties can significantly impact outcomes. A substantial portion of uncertainty in nuclear data originates from the subjective biases in the evaluation process, a crucial phase in the nuclear data production pipeline. Recent advancements indicate that automation of certain routines can mitigate these biases, thereby standardizing the evaluation process and enhancing reproducibility. This research aims to provide a methodology, framework, and metrics for the validation of automated nuclear data evaluation software leveraging high-quality synthetic data that closely mimic real experimental observables. An introduced error metric provides a scale and intuitive measure of the evaluation quality by quantifying the estimate’s accuracy and performance across the specified energy range. Synthetic data provides access to experimental observables and underlying resonance parameters, enabling comparison of different evaluations. The methodology is demonstrated using Ta-181 isotope data in the resolved resonance region. The Automated Resonance Identification Subroutine (ARIS), which operates without prior resonance information, was used to test and showcase the framework’s capabilities utilizing the proposed error metrics. The results demonstrate the effectiveness of the proposed approach and framework for optimizing software parameters and testing hypotheses through “what-if” controlled experiments, such as modifying assumptions about experimental conditions or average resonance parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-Driven State of Health Estimation for Second-Life Batteries Using Interpolated Synthetic Data and Feature Selection

Accurate estimation of the State of Health (SOH) for second-life batteries (SLBs) is crucial given their increasing use in energy storage applications. Precise SOH prediction is essential for safe operation and robust battery management systems. A major challenge is the limited availability of datasets for building reliable degradation models. To address this, synthetic data generation through linear interpolation is performed to extend the available data, making it more representative of real-world battery operating conditions. By analyzing feature correlation with SOH, the most relevant features are selected for the model. The proposed approach employs a convolutional neural network (CNN) model trained on this interpolated, feature-selected dataset, using time series data of voltage, temperature, and current over a cycle. By focusing on highly correlated features, the model achieves over 95% accuracy, with mean absolute error and root mean squared error up to 2.27% and 2.64%, respectively, in SOH estimation for two battery datasets tested. These results highlight the potential of combining synthetic data generation and feature selection to enhance SOH predictions, showcasing the superior performance of the proposed CNN model for both new batteries and SLBs.

feature selection↗

Creation Synthetic Data to Train a Digital Twin to Predict Reactor Operations

Understanding techniques to strengthen the nuclear safeguards regime is crucial in preventing nuclear proliferation due to recent advancements in the nuclear energy industry such as Generation IV reactors and microreactors. Prior to the construction of a nuclear power plant, it is necessary to understand the proliferation potential of the plant's reactor. Digital twins serve as a unique solution to recognizing reactor behavior indicative of nuclear proliferation. A digital twin is defined as a virtual model that works in unison to represent a physical asset, with a transference of data between the virtual and physical assets [1]. This work serves as validation for training a digital twin on synthetic data fabricated via means of Serpent reactor physics and point kinetics equations simulations. In this case this work is based on parameters of Idaho State University's AGN-201 reactor. The synthetic data can then be utilized to train machine learning models in the future to further investigate the utility of these methods. The accuracy of the predicted data is measured against real operational data to verify the reliability of the synthetic data creation methods and decide whether these methods should be used in the future to inform inspectors of a reactor's proliferation potential.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Data Efficiency Assessment of Generative Adversarial Networks for Critical Heat Flux Synthetic Data Generation

This study investigates the application of generative artificial intelligence techniques, particularly conditional generative adversarial networks (cGAN), in real-world engineering contexts, with a specific focus on synthetic data generation for critical heat flux (CHF). Utilizing a dataset comprising more than 20,000 real experimental CHF measurements, we conduct a series of experiments to examine cGAN’s behavior. These experiments encompass varying sizes of the training dataset, training cGAN on data from diverse experimental sources to generate new data on unseen experimental setups, and assessing the impact of excluding various input features on cGAN’s data generation accuracy. Our findings underscore the pronounced data dependency of cGAN for reliable performance, with decreased efficacy observed with smaller training dataset sizes. Notably, cGAN exhibits varying performance when trained on data from different experiments, with superior predictive capabilities observed for certain experiment sources compared to others. For instance, when cGAN was trained on data from Smolin et al.’s experiments or Zenkevich et al., it exhibited relatively good performance in generating the data from Becker et al., Kirillov et al., and Alekseev et al. experiments. In contrast, when trained with Alekseev et al.’s data and tasked with generating other experimental setups, cGAN showed notably poor performance. In both scenarios, cGAN’s performance was inferior compared to training on samples from all experiments concurrently. A feature importance analysis highlights the significant influence of parameters such as mass flux and heated length on accurate CHF generation, while other parameters like diameter and pressure have less impact. Inlet temperature is identified as a moderating factor by cGAN.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE↗

Utilizing Physics-Informed Synthetic Data to Train a Digital Twin for Predicting Reactor Operations

Understanding techniques to strengthen the nuclear nonproliferation regime is crucial in reducing the creation of nuclear weaponry on the basis of advancements in the nuclear energy industry. Prior to construction of a nuclear power plant, it is necessary to understand the proliferation potential of the plant’s reactor. Digital twins serve as a unique solution to recognizing reactor behavior indicative of nuclear proliferation. The following research conducted serves as a validation to training a digital twin on synthetic data fabricated via means of reactor physics simulations based on parameters of Idaho State University’s AGN-201 reactor. The synthetic data is utilized to train a long short-term memory (LSTM) recurrent neural network model. The accuracy of the predicted data is measured against real operational data to verify the reliability of the synthetic data creation methods and if these methods should be used in the future to inform inspectors of a reactor’s proliferation capabilities.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Applying Machine‐Learning Methods to Laser Acceleration of Protons: Lessons Learned From Synthetic Data

ABSTRACT In this study, we consider three different machine‐learning methods—a three‐hidden‐layer neural network, support vector regression, and Gaussian process regression—and compare how well they can learn from a synthetic data set for proton acceleration in the Target Normal Sheath Acceleration regime. The synthetic data set was generated from a previously published theoretical model by Fuchs et al. 2005 that we modified. Once trained, these machine‐learning methods can assist with efforts to maximize the peak proton energy, or with the more general problem of configuring the laser system to produce a proton energy spectrum with desired characteristics. In our study, we focus on both the accuracy of the machine‐learning methods and the performance on one GPU including memory consumption. Although it is arguably the least sophisticated machine‐learning model we considered, support vector regression performed very well in our tests.

Desai, Ronak↗

The Alcock–Paczynski effect from Lyman- α forest correlations: analysis validation with synthetic data

The three-dimensional distribution of the Ly α forest has been extensively used to constrain cosmology through measurements of the baryon acoustic oscillations (BAO) scale. However, more cosmological information could be extracted from the full shapes of the Ly α forest correlations through the Alcock–Paczynski (AP) effect. In this work, we prepare for a cosmological analysis of the full shape of the Ly α forest correlations by studying synthetic data of the extended Baryon Oscillation Spectroscopic Survey (eBOSS). We use a set of 100 eBOSS synthetic data sets in order to validate such an analysis. These mocks undergo the same analysis process as the real data. We perform a full-shape analysis on the mean of the correlation functions measured from the 100 eBOSS realizations, and find that our model of the Ly α correlations performs well on current data sets. We show that we are able to obtain an unbiased full-shape measurement of D M /D H (z eff ), where D M is the transverse comoving distance, D H is the Hubble distance, and z eff is the effective redshift of the measurement. We test the fit over a range of scales, and decide to use a minimum separation of r min = 25 h –1 Mpc. Here, we also study and discuss the impact of the main contaminants affecting Ly α forest correlations, and give recommendations on how to perform such analysis with real data. While the final eBOSS Ly α BAO analysis measured D M /D H (z eff = 2.33) with 4 per cent statistical precision, a full-shape fit of the same correlations could provide an $\sim 2~{{\ \rm per\ cent}}$ measurement.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerated alloy discovery using synthetic data generation and data mining

The search for new alloys with improved properties is never ending with infinite combinations and amounts of alloying elements in the alloy. Advancements in machine learning have made navigating this enormous search space feasible. However, training the machine learning models and tuning their hyper-parameters to make accurate predictions can be time-consuming and often require high-performance computing resources. Furthermore, the quality of the predictions depend on the availability of sufficient training data. Here, we present a generic approach to accelerate alloy discovery by coupling high throughput CALPHAD calculations, synthetic data generation, and data mining. Finally, as a demonstration of the approach, we design super bainitic steels that form bainite at 200 C in lower transformation times.

36 MATERIALS SCIENCE↗

Models for synthetic data generation

The software includes a suite of probabilistic statistical/machine learning models that can generate discrete synthetic data. Each model is trained on a set of real (private) data and then it can be used to generate synthetic but statistically similar data. Once ready, the model can generate as many samples as we want. Finally, in addition to the actual models, the software includes code to process data, evaluate results (based on cross validation), and produce reports.

De Oliveira Sales, Ana Paula↗

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A machine learning estimator trained on synthetic data for real-time earthquake ground-shaking predictions in Southern California

Abstract After large-magnitude earthquakes, a crucial task for impact assessment is to rapidly and accurately estimate the ground shaking in the affected region. To satisfy real-time constraints, intensity measures are traditionally evaluated with empirical Ground Motion Models that can drastically limit the accuracy of the estimated values. As an alternative, here we present Machine Learning strategies trained on physics-based simulations that require similar evaluation times. We trained and validated the proposed Machine Learning-based Estimator for ground shaking maps with one of the largest existing datasets (<100M simulated seismograms) from CyberShake developed by the Southern California Earthquake Center covering the Los Angeles basin. For a well-tailored synthetic database, our predictions outperform empirical Ground Motion Models provided that the events considered are compatible with the training data. Using the proposed strategy we show significant error reductions not only for synthetic, but also for five real historical earthquakes, relative to empirical Ground Motion Models.

Environmental Sciences & Ecology↗