Search NASASearch

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis

Synthetic Hyperspectral Data for Global Water Quality Algorithm Development

Eutrophication and increasing prevalence of potentially toxic algal blooms (cyanoHABs) among global inland water bodies have become a major ecological concern and require direct attention. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local aquatic processes, to spatially resolved global products. Planned aquatic biogeochemistry remote sensing data products from hyperspectral imagers such as NASA’s Surface Biology and Geology (SBG) mission and relevant aquatic sensor sensitivity precursor airborne imaging spectrometer data provide unprecedented radiometric resolution and sensor sensitivity for characterizing complex aquatic ecosystems. However, scarcity of high-quality freshwater in-situ optical data hinders our capability to develop and validate robust retrieval algorithms. A state-of-the-art synthetic dataset of paired top-of-atmosphere, bottom-of-atmosphere, and optical and biogeophysical data was developed through radiative transfer modeling to simulate natural freshwater ecosystems. A synthetic or precursor dataset for SBG is being used to train robust machine learning models to derive water quality products pertinent to SBG mission objectives. The dataset is also used to show the potential of performing vigorous aquatic sensitivity studies and explored pathways for how best to optimize hyperspectral data for machine learning development. A processing pipeline and resultant global synthetic/precursor dataset for inland waters is presented to establish the innovation for water quality studies of inland waters globally. Optical Society of America Imaging and Applied Optics Congress, Hyperspectral Imaging and Sounding of the Environment (OSA HISE) Meeting, 19-23 July 2021, Virtual Meeting, https://www.osa.org/enus/meetings/osa_meetings/optical_sensors_and_sensing_congress/program/hyperspectral_imaging_and_sounding_of_the_environm/

Synthetic

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]

Inversion of limb radiance measurements - An operational algorithm

The limb radiance inversion radiometer (LRIR) and limb infrared monitor of the stratosphere (LIMS) experiments aboard the Nimbus 6 and 7 spacecraft have made observations of infrared emission by CO2, O3, H2O, HNO3, and NO2 at the earth's limb. This paper describes a method by which such measurements can be inverted to give vertical distributions of temperature and mixing ratios as functions of pressure. The simple and efficient approach was successfully applied to the LRIR data and subsequently in the initial assessment of the LIMS data. Inversion of synthetic data indicates the size of the errors to be expected as a result of the assumptions and instrumental errors. Retrievals of measured LIMS radiances are shown as examples and compared to in situ observations. The differences are comparable to those obtained with the more complex retrieval scheme used to process the LIMS archival products. Some problems are noted.

Bailey, P. L.

Helioseismic Measurements of Convective Power in Solar Cycle 24

Constraining the parameters under which convection in the solar interior operates has important implications for describing how energy is transported by plasma motions, and various physical models have been employed to provide some theoretical estimates on the expected power spectrum. Past attempts to measure the convective power distributed among large spatial scales have, however, found differing and incompatible values. Here, we present measurements of the convective power spectrum in the upper convection zone for Carrington rotations in Solar Cycle 24 obtained from the helioseismic signal corresponding to East-West flows. We also perform calibration on synthetic data using the global acoustic GALE code to make assessments of the flow velocities at various length scales without the need for performing inversions. This allows us to derive the convective power spectrum from the flow maps and to compare with the results from the raw travel times. These results are compared against predictions of convective power in global models of convection produced by the EULAG code. Finally, we show how the steps in our analysis procedure (for example data segmentation, filtering, etc.) affect our estimates of the convective power by comparing with the synthetic data from the GALE code.

Heliophysics

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil

Estimating the Single-Trial Characteristics of Event-Related Responses: Evaluation of the MCERP Algorithm

Single-trial event-related responses collected during the course of an experiment are typically averaged before analysis resulting in a rather crude picture of event-related brain dynamics. It has been quite clear for some time that these responses exhibit trial-to-trial variability: however, the computational techniques necessary to deal with such responses in noisy conditions have not been available. To this end we have developed the multiple-component, event-related potential model (mcERP), which assumes that the each event-related response consists of a sum of multiple evoked components each described by a stereotypical waveshape. These waveshapes are allowed to vary in amplitude and onset latency from trial to trial, which allows us to capture, to first-order, the trial-dependent variations in event-related brain dynamics. We have constructed many sets of synthetic data designed to simulate intracortical recordings from a 15 channel, linear-array multielectrode implanted acutely in V1 of an awake-behaving macaque undergoing visual stimulation with a red light flash. This synthetic data was used to characterize the performance of the mcERP algorithm. First we quantified the degree to which such trial-to-trial variability aids in the identification of multiple components, and we demonstrate that amplitude variability is a more important factor in component separation than latency variability. Second, we quantified the behavior of the algorithm under two distinct signal-to-noise ratio (SNR) conditions: Gaussian noise independently present in each channel, and highly correlated (1/f distributed), far-field noise presented identically in each channel of the array. The mcERP algorithm was found to be robust to noise accurately identifying all component waveshapes and their associated single-trial characteristics down to SNR levels of -20dB for Gaussian noise and -7dB for 1/f far-field noise. Comparisons of the performance of this algorithm with factor analysis (FA) and independent component analysis (ICA) will be described by Knuth et al. (SFN abstracts, 2002). In addition, the advantages of application of mcERP to real data will be described by Shah et al, (these abstracts, 2002: SFN abstracts, 2002).

Knuth, K. H.

Maven: a multimodal foundation model for supernova science

Abstract A common setting in astronomy is the availability of a small number of high-quality observations, and larger amounts of either lower-quality observations or synthetic data from simplified models. Time-domain astrophysics is a canonical example of this imbalance, with the number of supernovae observed photometrically outpacing the number observed spectroscopically by multiple orders of magnitude. At the same time, no data-driven models exist to understand these photometric and spectroscopic observables in a common context. Contrastive learning objectives, which have grown in popularity for aligning distinct data modalities in a shared embedding space, provide a potential solution to extract information from these modalities. We present Maven, the first foundation model for supernova science. To construct Maven, we first pre-train our model to align photometry and spectroscopy from 0.5 M synthetic supernovae using a contrastive objective. We then fine-tune the model on 4702 observed supernovae from the Zwicky transient facility. Maven reaches state-of-the-art performance on both classification and redshift estimation, despite the embeddings not being explicitly optimized for these tasks. Through ablation studies, we show that pre-training with synthetic data improves overall performance. In the upcoming era of the Vera C. Rubin observatory, Maven will serve as a valuable tool for leveraging large, unlabeled and multimodal time-domain datasets.

Zhang, Gemma (ORCID:0000000280198082)

Terrestrial Water Mass Load Changes from Gravity Recovery and Climate Experiment (GRACE)

Recent studies show that data from the Gravity Recovery and Climate Experiment (GRACE) is promising for basin- to global-scale water cycle research. This study provides varied assessments of errors associated with GRACE water storage estimates. Thirteen monthly GRACE gravity solutions from August 2002 to December 2004 are examined, along with synthesized GRACE gravity fields for the same period that incorporate simulated errors. The synthetic GRACE fields are calculated using numerical climate models and GRACE internal error estimates. We consider the influence of measurement noise, spatial leakage error, and atmospheric and ocean dealiasing (AOD) model error as the major contributors to the error budget. Leakage error arises from the limited range of GRACE spherical harmonics not corrupted by noise. AOD model error is due to imperfect correction for atmosphere and ocean mass redistribution applied during GRACE processing. Four methods of forming water storage estimates from GRACE spherical harmonics (four different basin filters) are applied to both GRACE and synthetic data. Two basin filters use Gaussian smoothing, and the other two are dynamic basin filters which use knowledge of geographical locations where water storage variations are expected. Global maps of measurement noise, leakage error, and AOD model errors are estimated for each basin filter. Dynamic basin filters yield the smallest errors and highest signal-to-noise ratio. Within 12 selected basins, GRACE and synthetic data show similar amplitudes of water storage change. Using 53 river basins, covering most of Earth's land surface excluding Antarctica and Greenland, we document how error changes with basin size, latitude, and shape. Leakage error is most affected by basin size and latitude, and AOD model error is most dependent on basin latitude.

Seo, K.-W.

Application of machine learning techniques for fast MeV x-ray spectra unfolding from filter stack spectrometer data

Recovery of MeV x-ray spectra from detector signals is difficult because the response matrix inversion is ill-conditioned and current methods are too slow for high-repetition-rate experiments. In this work, we make use of neural networks to unfold MeV x-ray spectra from measurements obtained with a filter stack spectrometer at rates of near 40 Hz. The neural network was trained on synthetic data and tested on both synthetic and experimental data, the latter obtained in two separate experiments performed at the Omega EP laser facility. We show here that this unfolding method has good performance on synthetic data and that it is a promising option for experimental data of up to 40 MeV. The accuracy on experimental data is verified by using a simple forward model to compare against measured values.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Bayesian Optimization of Non-Invariant Systems with Constraints Developed for Application to the ECR Ion Source VENUS

In this work, we consider the optimization of non-invariant systems with both safety and control constraints. We present a new approach based on Bayesian optimization for the dynamic, safe and controlled optimization of such systems. Although there are other possible use cases, we focus on the application to the electron cyclotron resonance ion source VENUS. From experimental data, we have observed that VENUS behaves to first order as a non-invariant dynamic system with moving areas of instability. Our novel approach aims at providing a tool that can maintain system optimization in a safe way. This is accomplished by making sure the objective function, the beam current in the case of VENUS, does not fall under an operational minimum, while simultaneously requiring the optimization to avoid areas where VENUS is unstable. We compare the result of our approach on synthetic data modeled to mimic the behavior of VENUS with two methods from the literature, a standard Bayesian optimizer and a safe Bayesian optimizer, both adapted to deal with dynamic systems. A cross Student T-test is conducted to show the significance of the improvement given by the new method we introduce here, regarding the two preexisting methods we compared to. The results of the tests conducted on synthetic data show that the proposed method succeeds at maintaining the system optimized and obeys the predefined constraints better than the literature methods explored.

Bayesian optimization

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Machine learning analysis of high-repetition-rate two-dimensional Thomson scattering spectra from laser-produced plasmas

With the emergence of high-repetition-rate two-dimensional Thomson scattering (TS) measurements, improving spectral data analysis is a key area of interest. Here, we present a new way to derive the electron temperature and density of laser-driven blast waves in plasmas from their TS spectra with machine learning (ML). This analysis occurs in both the non-collective (α < 1) and collective (α > 1) scattering regimes with the goal of autonomously and more accurately determining T c and n e both where spectral data has been collected and to give the ability to predict these attributes in regions where data has not been collected. We introduce three ML models, one trained only on experimental data, one only on synthetic data, and one using transfer learning, and compare their speed and accuracy with the conventional TS inversion algorithms in the open source PlasmaPy python package.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

DRDMannTurb: A Python package for scalable, data-driven synthetic turbulence

Synthetic turbulence models (STMs) are used in wind engineering to generate realistic flow fields and are employed as inputs to industrial wind simulations. Examples include prescribing inlet conditions in large eddy simulations that model loads on wind turbines and tall buildings. We are interested in STMs capable of generating fluctuations based on prescribed second-moment statistics since such models can simulate environmental conditions that closely resemble on-site observations. To this end, the widely used Mann model (see Mann, 1994, 1998) is the inspiration for DRDMannTurb. The Mann model is described by three physical parameters: a magnitude parameter influencing the global variance of the wind field and corresponding to the Kolmogorov constant multiplied by the rate of viscous dissipation of the turbulent kinetic energy to the two-thirds, αϵ 2/3 , a turbulence length scale parameter L, and a nondimensional parameter Γ related to the lifetime of the eddies. A number of studies, as well as international standards (e.g., those by the International Electrotechnical Commission (IEC)), include recommended values for these three parameters with the goal of standardizing wind simulations according to observed energy spectra. Yet, having only three parameters, the Mann model faces limitations in accurately representing the diversity of observable spectra. This Python package enables users to extend the Mann model and more accurately fit field measurements through flexible neural network models of the eddy lifetime function. Following Keith et al. (2021), we refer to this class of models as Deep Rapid Distortion (DRD) models. DRDMannTurb also includes a general module implementing an efficient method for synthetic turbulence generation based on a domain decomposition technique. This technique is also described in Keith et al. (2021).

17 WIND ENERGY