Search NASASearch

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Validating automated resonance evaluation with synthetic data

The integrity and precision of nuclear data are crucial for a broad spectrum of applications, from national security and nuclear reactor design to medical diagnostics, where the associated uncertainties can significantly impact outcomes. A substantial portion of uncertainty in nuclear data originates from the subjective biases in the evaluation process, a crucial phase in the nuclear data production pipeline. Recent advancements indicate that automation of certain routines can mitigate these biases, thereby standardizing the evaluation process and enhancing reproducibility. This research aims to provide a methodology, framework, and metrics for the validation of automated nuclear data evaluation software leveraging high-quality synthetic data that closely mimic real experimental observables. An introduced error metric provides a scale and intuitive measure of the evaluation quality by quantifying the estimate’s accuracy and performance across the specified energy range. Synthetic data provides access to experimental observables and underlying resonance parameters, enabling comparison of different evaluations. The methodology is demonstrated using Ta-181 isotope data in the resolved resonance region. The Automated Resonance Identification Subroutine (ARIS), which operates without prior resonance information, was used to test and showcase the framework’s capabilities utilizing the proposed error metrics. The results demonstrate the effectiveness of the proposed approach and framework for optimizing software parameters and testing hypotheses through “what-if” controlled experiments, such as modifying assumptions about experimental conditions or average resonance parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Backus Effect on a Perpendicular Errors in Harmonic Models of Real vs. Synthetic Data

Measurements of geomagnetic scalar intensity on a thin spherical shell alone are not enough to separate internal from external source fields; moreover, such scalar data are not enough for accurate modeling of the vector field from internal sources because of unmodeled fields and small data errors. Spherical harmonic models of the geomagnetic potential fitted to scalar data alone therefore suffer from well-understood Backus effect and perpendicular errors. Curiously, errors in some models of simulated 'data' are very much less than those in models of real data. We analyze select Magsat vector and scalar measurements separately to illustrate Backus effect and perpendicular errors in models of real scalar data. By using a model to synthesize 'data' at the observation points, and by adding various types of 'noise', we illustrate such errors in models of synthetic 'data'. Perpendicular errors prove quite sensitive to the maximum degree in the spherical harmonic expansion of the potential field model fitted to the scalar data. Small errors in models of synthetic 'data' are found to be an artifact of matched truncation levels. For example, consider scalar synthetic 'data' computed from a degree 14 model. A degree 14 model fitted to such synthetic 'data' yields negligible error, but amplifies 4 nT (rmss) added noise into a 60 nT error (rmss); however, a degree 12 model fitted to the noisy 'data' suffers a 492 nT error (rmms through degree 12). Geomagnetic measurements remain unaware of model truncation, so the small errors indicated by some simulations cannot be realized in practice. Errors in models fitted to scalar data alone approach 1000 nT (rmss) and several thousand nT (maximum).

Voorhies, C. V.

Data-Driven State of Health Estimation for Second-Life Batteries Using Interpolated Synthetic Data and Feature Selection

Accurate estimation of the State of Health (SOH) for second-life batteries (SLBs) is crucial given their increasing use in energy storage applications. Precise SOH prediction is essential for safe operation and robust battery management systems. A major challenge is the limited availability of datasets for building reliable degradation models. To address this, synthetic data generation through linear interpolation is performed to extend the available data, making it more representative of real-world battery operating conditions. By analyzing feature correlation with SOH, the most relevant features are selected for the model. The proposed approach employs a convolutional neural network (CNN) model trained on this interpolated, feature-selected dataset, using time series data of voltage, temperature, and current over a cycle. By focusing on highly correlated features, the model achieves over 95% accuracy, with mean absolute error and root mean squared error up to 2.27% and 2.64%, respectively, in SOH estimation for two battery datasets tested. These results highlight the potential of combining synthetic data generation and feature selection to enhance SOH predictions, showcasing the superior performance of the proposed CNN model for both new batteries and SLBs.

feature selection

Data Efficiency Assessment of Generative Adversarial Networks for Critical Heat Flux Synthetic Data Generation

This study investigates the application of generative artificial intelligence techniques, particularly conditional generative adversarial networks (cGAN), in real-world engineering contexts, with a specific focus on synthetic data generation for critical heat flux (CHF). Utilizing a dataset comprising more than 20,000 real experimental CHF measurements, we conduct a series of experiments to examine cGAN’s behavior. These experiments encompass varying sizes of the training dataset, training cGAN on data from diverse experimental sources to generate new data on unseen experimental setups, and assessing the impact of excluding various input features on cGAN’s data generation accuracy. Our findings underscore the pronounced data dependency of cGAN for reliable performance, with decreased efficacy observed with smaller training dataset sizes. Notably, cGAN exhibits varying performance when trained on data from different experiments, with superior predictive capabilities observed for certain experiment sources compared to others. For instance, when cGAN was trained on data from Smolin et al.’s experiments or Zenkevich et al., it exhibited relatively good performance in generating the data from Becker et al., Kirillov et al., and Alekseev et al. experiments. In contrast, when trained with Alekseev et al.’s data and tasked with generating other experimental setups, cGAN showed notably poor performance. In both scenarios, cGAN’s performance was inferior compared to training on samples from all experiments concurrently. A feature importance analysis highlights the significant influence of parameters such as mass flux and heated length on accurate CHF generation, while other parameters like diameter and pressure have less impact. Inlet temperature is identified as a moderating factor by cGAN.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE

Quantifying Errors in 3D CME Parameters Derived from Synthetic Data Using White-Light Reconstruction Techniques

Current efforts in space weather forecasting of CMEs have been focused on predicting their arrival time and magnetic structure. To make these predictions, methods have been developed to derive the true CME speed, size, position, and mass, among others. Difficulties in determining the input parameters for CME forecasting models arise from the lack of direct measurements of the coronal magnetic fields and uncertainties in estimating the CME 3D geometric and kinematic parameters after eruption. White-light coronagraph images are usually employed by a variety of CME reconstruction techniques that assume more or less complex geometries. This is the first study from our International Space Science Institute (ISSI) team “Understanding Our Capabilities in Observing and Modeling Coronal Mass Ejections”, in which we explore how subjectivity affects the 3D CME parameters that are obtained from the Graduated Cylindrical Shell (GCS) reconstruction technique, which is widely used in CME research. To be able to quantify such uncertainties, the “true” values that are being fitted should be known, which are impossible to derive from observational data. We have designed two different synthetic scenarios where the “true” geometric parameters are known in order to quantify such uncertainties for the first time. We explore this by using two sets of synthetic data: 1) Using the ray-tracing option from the GCS model software itself, and 2) Using 3D magnetohydrodynamic (MHD) simulation data from the Magnetohydrodynamic Algorithm outside a Sphere code. Our experiment includes different viewing configurations using single and multiple viewpoints. CME reconstructions using a single viewpoint had the largest errors and error ranges overall for both synthetic GCS and simulated MHD white-light data. As the number of viewpoints increased from one to two, the errors decreased by approximately 4° in latitude, 22° in longitude, 14° in tilt, and 10° in half-angle. Our results quantitatively show the critical need for at least two viewpoints to be able to reduce the uncertainty in deriving CME parameters. We did not find a significant decrease in errors when going from two to three viewpoints for our specific hypothetical three spacecraft scenario using synthetic GCS white-light data. As we expected, considering all configurations and numbers of viewpoints, the mean absolute errors in the measured CME parameters are generally significantly higher in the case of the simulated MHD white-light data compared to those from the synthetic white-light images generated by the GCS model. We found the following CME parameter error bars as a starting point for quantifying the minimum error in CME parameters from white-light reconstructions: Δθ (latitude)=6° +2° -3° , Δϕ (longitude)=11° +18° -6° , Δγ (tilt)=25° +8° -7° , Δx (half-angle)=10° +12° -6° , Δh (height)=0.6 +1.2 -0.4 R ⨀ , and Δκ (ratio)=0.1 +0.03 -0.02 .

Coronal mass ejections

Quantifying Errors in 3D CME Parameters Derived From Synthetic Data Using White-Light Reconstruction Techniques

Current efforts in space weather forecasting of CMEs have been focused on predicting their arrival time and magnetic structure. To make these predictions, methods have been developed to derive the true CME speed, size, position, and mass, among others. Difficulties in determining the input parameters for CME forecasting models arise from the lack of direct measurements of the coronal magnetic fields and uncertainties in estimating the CME 3D geometric and kinematic parameters after eruption. White-light coronagraph images are usually employed by a variety of CME reconstruction techniques that assume more or less complex geometries. This is the first study from our International Space Science Institute (ISSI) team “Understanding Our Capabilities in Observing and Modeling Coronal Mass Ejections”, in which we explore how subjectivity affects the 3D CME parameters that are obtained from the Graduated Cylindrical Shell (GCS) reconstruction technique, which is widely used in CME research. To be able to quantify such uncertainties, the “true” values that are being fitted should be known, which are impossible to derive from observational data. We have designed two different synthetic scenarios where the “true” geometric parameters are known in order to quantify such uncertainties for the first time. We explore this by using two sets of synthetic data: 1) Using the ray-tracing option from the GCS model software itself, and 2) Using 3D magnetohydrodynamic (MHD) simulation data from the Magnetohydrodynamic Algorithm outside a Sphere code. Our experiment includes different viewing configurations using single and multiple viewpoints. CME reconstructions using a single viewpoint had the largest errors and error ranges overall for both synthetic GCS and simulated MHD white-light data. As the number of viewpoints increased from one to two, the errors decreased by approximately 4° in latitude, 22° in longitude, 14° in tilt, and 10° in half-angle. Our results quantitatively show the critical need for at least two viewpoints to be able to reduce the uncertainty in deriving CME parameters. We did not find a significant decrease in errors when going from two to three viewpoints for our specific hypothetical three spacecraft scenario using synthetic GCS white-light data. As we expected, considering all configurations and numbers of viewpoints, the mean absolute errors in the measured CME parameters are generally significantly higher in the case of the simulated MHD white-light data compared to those from the synthetic white-light images generated by the GCS model. We found the following CME parameter error bars as a starting point for quantifying the minimum error in CME parameters from white-light reconstructions: Δθ (latitude)=6° +2° -3° , Δϕ (longitude)=11° +18° -6° , Δγ (tilt)=25° +8° -7° , Δx (half-angle)=10° +12° -6° , Δh (height)=0.6 +1.2 -0.4 R ⨀ , and Δκ (ratio)=0.1 +0.03 -0.02 .

Coronal mass ejections

Applying Machine‐Learning Methods to Laser Acceleration of Protons: Lessons Learned From Synthetic Data

ABSTRACT In this study, we consider three different machine‐learning methods—a three‐hidden‐layer neural network, support vector regression, and Gaussian process regression—and compare how well they can learn from a synthetic data set for proton acceleration in the Target Normal Sheath Acceleration regime. The synthetic data set was generated from a previously published theoretical model by Fuchs et al. 2005 that we modified. Once trained, these machine‐learning methods can assist with efforts to maximize the peak proton energy, or with the more general problem of configuring the laser system to produce a proton energy spectrum with desired characteristics. In our study, we focus on both the accuracy of the machine‐learning methods and the performance on one GPU including memory consumption. Although it is arguably the least sophisticated machine‐learning model we considered, support vector regression performed very well in our tests.

Desai, Ronak

Synthetic Data Generation for 3D Mesh Prediction and Spatial Reasoning During Multi-Agent Robotic Missions

In-space assembly operations require accurate reasoning over the pose, location, and structural organization of both the autonomous agents and assembly materials. In a full six-degree-of-freedom space, an accurate understanding of the full three-dimensional structure of the object of interest greatly enriches information for pose estimation and collision planning. Current methods of predicting pose estimation require a priori understanding of the shape of the object. Additionally, visual information in the space environment is impacted by variations in contrast and illumination. Using synthetic data allows us to rapidly generate large datasets with in varying environments and lighting conditions.This work details the generation of synthetic data used to explore the use of a region-based convolutional neural networks to detect objects of interest and predict a voxel-based three-dimensional mesh in order to understand their full three-dimensional shape. This mesh provides useful spatial information during in-space assembly operations without requiring either the complexity of maintaining models over the progress of building an object or observations from multiple angles. The generated meshes are then compared to that of ground truth in order to measure its performance.

synthetic data

Synthetic Data Generation for 3D Mesh Prediction and Spatial Reasoning During Multi-Agent Robotic Missions

In-space assembly operations require accurate reasoning over the pose, location, and structural organization of both the autonomous agents and assembly materials. In a full six-degree-of-freedom space, an accurate understanding of the full three-dimensional structure of the object of interest greatly enriches information for pose estimation and collision planning. Current methods of predicting pose estimation require a priori understanding of the shape of the object. Additionally, visual information in the space environment is impacted by variations in contrast and illumination. Using synthetic data allows us to rapidly generate large datasets with in varying environments and lighting conditions. This work details the generation of synthetic data used to explore the use of a region-based convolutional neural networks to detect objects of interest and predict a voxel-based three-dimensional mesh in order to understand their full three-dimensional shape. This mesh provides useful spatial information during in-space assembly operations without requiring either the complexity of maintaining models over the progress of building an object or observations from multiple angles. The generated meshes are then compared to that of ground truth in order to measure its performance.

James Ecker

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING

Development of the Ames Global Hyperspectral Synthetic Data Set: Surface Bidirectional Reflectance Distribution Function

This study introduces the Ames Global Hyperspectral Synthetic Data set (AGHSD), in particular the surface bidirectional reflectance distribution function (BRDF) product, to support the NASA Surface Biology and Geology (SBG) mission development. The data set is generated based on the corresponding multispectral BRDF products from NASA's MODIS satellite sensor. Based on theories of radiative transfer in vegetation canopies, we derive a simple but robust relationship that indicates that the hyperspectral surface BRDF can be accurately approximated as a weighted sum of the soil surface reflectance, the leaf single albedo, and the canopy scattering coefficient, where the weights or coefficients are spectrally invariant and thus readily estimated from the multispectral MODIS products. We validate the algorithm with simulations by a Monte Carlo Ray Tracing model and find the results highly consistent with the theoretic derivation. Using reflectance spectra of soil and vegetation derived from existing spectral libraries, we apply the algorithm to generate the AGHSD BRDF product at 1 km and 8-day resolutions for the year of 2019. The data set is biogeochemically and biogeophysically coherent and consistent, and serves the goal to support the SBG community in developing sciences and applications for the future global imaging spectroscopy mission.

Hyperspectral Remote Sensing

Updates in Developing a Prototype Science Pipeline and Full-Volume, Global Hyperspectral Synthetic Data Sets for NASA’s Earth System Observatory’s Upcoming Surface, Biology and Geology Mission

The Surface Biology and Geology (SBG) mission recently passed mission confirmation review and has entered phase A – design and development. SBG will acquire high resolution solar-reflected spectroscopy and thermal infrared observations at a data rate of ~2.5 TB/day and generate products at ~40 TB/day. Given that the per-day volume is greater than NASA’s total extant airborne hyperspectral data collection, collecting, processing, disseminating, and exploiting the SBG data present new challenges. To meet these challenges, we have developed a prototype science pipeline and a full-volume global hyperspectral synthetic data set to help prepare for SBG’s flight (see poster GC42D-0730). Our science pipeline is based on the science processing technology developed for NASA’s Kepler and TESS planet-hunting missions. The pipeline infrastructure, Ziggy, provides a scalable architecture for robust, repeatable, and replicable science and application products that can be run on a range of systems from a laptop to the cloud or a supercomputer. Ziggy is compliant with NASA Procedural Requirement (NPR) 7150.2C, is at a technical readiness level (TRL) of 7 and has been released to github.com/nasa/ziggy. We integrated Ziggy with EO-1/Hyperion workflows to build a prototype pipeline and ingested the 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. We fully implemented the first stage and processed the entire 55 TB Hyperion data set from the raw data (Level 0) to top-of-the-atmosphere radiance (Level 1R). We are currently evaluating the ISOFIT atmospheric correction module to convert the L1R data to surface reflectance (Level 2) before reprocessing the full data set to L2. Crosschecks are being performed with RadCalNet as well as with coincident observations by AVIRIS. We are also investigating modern methods for georectifying the Hyperion scenes. Finally, we describe an analysis of the cost to conduct forward processing and reprocessing campaigns for SBG on HECC with dedicated compute and storage resources using the resurrected Hyperion pipeline as a proxy for full-volume SBG data. The analysis demonstrates that SBG L0 data can be processed to L2 on HECC with full reprocessing campaigns every two years for ~$2.6M over a 7-year lifespan. Moreover, 69% of the system capacity would be available for other activities, possibly enabling future open-source science activities, including algorithm development, L3+ processing, .etc.

ESD

Synthetic Data for Testing TRMM Radar Algorithms

Test data are required to test algorithms for the TRMM Precipitation Radar. These data are needed to test the design of the computer codes under development for the operational phase of the mission, and also to test and evaluate alternative or improved precipitation retrieval algorithms. Over a number of years we have developed and used a 3-dimensional radar model for simulating spaceborne precipitation radars. We have adapted this code to produce data files as close as possible to the TRMM file specifications. In this paper, we will describe the model as it is currently implemented, and show some samples of the synthetic data sets.

Jones, Jeffrey A.

Small target detection for search and rescue operations using distributed deep learning and synthetic data generation

It is important to find the target as soon as possible for search and rescue operations. Surveillance camera systems and unmanned aerial vehicles (UAVs) are used to support search and rescue. Automatic object detection is important because a person cannot monitor multiple surveillance screens simultaneously for 24 hours. Also, the object is often too small to be recognized by the human eye on the surveillance screen. This study used UAVs around the Port of Houston and fixed surveillance cameras to build an automatic target detection system that supports the US Coast Guard (USCG) to help find targets (e.g., person overboard). We combined image segmentation, enhancement, and convolution neural networks to reduce detection time to detect small targets. We compared the performance between the auto-detection system and the human eye. Our system detected the target within 8 seconds, but the human eye detected the target within 25 seconds. Our systems also used synthetic data generation and data augmentation techniques to improve target detection accuracy. This solution may help the search and rescue operations of the first responders in a timely manner.

Chow, Edward

An Investigation Into Possible Systematic Effects on Neutron Star Radius Estimates using NICER-like Synthetic Data

Neutron star cores contain the densest matter in the observable universe. The state of this matter is of interest in numerous fields, but laboratory experiments cannot explore this matter. Although the composition of the matter would be of great interest, macroscopic observables such as the neutron star mass-radius relation depend primarily on the equation of state (EOS). As a result, precise and reliable radius measurements would be valuable in constraining the EOS. However, most attempts at radius measurements are susceptible to systematic errors, meaning the inferred radius can be significantly biased even though the fit to the data appears to be statistically good. Previous studies suggested that radii inferred using X-ray data provided by NASA’s NICER mission may be more immune from such systematic errors. This is in part because, compared with previous measurements that obtained averaged spectra and fluxes, NICER observes millisecond pulsars by timing the photons so precisely, it is possible to obtain the spectrum as a function of rotational phase and see variations such as heated regions on the star rotate into and out of view hundreds of times per second. This extra information seems promising to break degeneracies and mitigate systematic errors, but a more in-depth study is necessary. We report the first steps of that study, in which we generate NICER-like synthetic data and determine the quality of fit and bias in radius obtained when we fit the data using a model different from the model used to generate the data.

Isiah Holt