Synthetic Data Generation Using GenAI-Based WGAN
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.
An attempt is made to improve the simulation of the meso-alpha and meso-beta scales' precipitation distribution over the Florida peninsula during the warm season through improvement of the initialization of beta scale convective cloud systems, using GOES IR and visible satellite imagery. These data can be used to define the location and spatial extent of meso-beta scale cloud systems at a specified time. Attention is presently given to the results of the first phase of this study, in which only moisture information was synthesized from the satellite imagery.
The inversion of electromagnetic information to physical and biological properties of the water column is a notoriously difficult problem, yet fundamental to our ability of understanding aquatic processes on large time and space scales. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will discuss research integrating advanced biological and radiative modeling, high-end computation, and machine learning to develop a portable global processor for simultaneous retrieval of atmosphere and water optics for diverse aquatic systems from the open and coastal ocean to optically extreme inland waters and harmful algal blooms. We will discuss some of the basic concepts behind the forward modeling approach including DEAP, the novel Distributed Equivalent Algal Populations model, for developing large spectral libraries of aquatic particle optics to aid in our ability to distinguish phytoplankton functional types (PFTs) and inorganic material, as well as other factors which enable comprehensive modeling from the benthos to top-of-atmosphere (TOA). This information is being used to understand how we can leverage next-generation deep learning methods for maximum information retrieval and rapid image processing, while also providing capabilities to identify minimum sensor spectral requirements necessary for certain aquatic applications. Further, I will touch on how we envision this research to enable the aquatic community for science discovery and how we are moving closer towards the capability for high-fidelity global analysis of aquatic ecosystems.
Ring diagrams are cross-sections of three-dimensional spatiotemporal power spectra of solar oscillations. The rings reveal information about sub-surface flows and represent an important tool for helioseismology. Ring diagrams can be constructed using Doppler velocity or intensity maps of spectral lines. How the velocities are computed is an important factor for accuracy of information we can retrieve on subsurface flows. In our work, the ring diagrams are generated from Doppler shift data of synthesized Fe I 6173 Å line. We compare ring diagrams computed by two methods–HMI line-of-sight pipeline and the bisector of Fe line. Fe I line is synthesized for StellarBox 3D Radiative hydrodynamic simulations under LTE assumption. We aim to answer the following questions: 1.How do power spectra obtained from velocities computed with the HMI pipeline and bisector of the Fe I 6173Å compare? 2.How do the power spectral density retrieved with each method vary with heliocentric angle? 3.What is the effect of changing resolution on the power spectral density in ring diagrams obtained with the two methods?
Synthetic aperture radar (SAR) has high data rate because it collects and processes the data coherently. The data rate limitation of the system has to be satisfied while maintaining good image quality. Thus, a quantizer with minimum data rate and high SNR should be employed. An adaptive quantization method is proposed for the burst mode SAR. This adaptive quantizer uses uniformly quantized data to select a subset of bits which is equivalent to changing the step size of the uniform quantizer. A simple implementation which uses the previous burst data to compute the local statistics for the bit selection is presented. The use of previous burst simplifies the implementation because it does not require storage or delay; however, an abrupt change in the terrain could result in incorrect bit selection. An error analysis of this implementation and comparison of two burst mode SAR images formed using the uniformly quantized and adaptively quantized data is presented.
Synthetic aperture radars (SAR) image from a non-nadir position. Thus the orientation of the target and sensor to one another is of paramount importance. This has posed problems for data interpretation and with the potentials of radar data for change detection studies. It is possible to use Seasat radar data for change detection even though the look directions are fixed for each location. Especially in areas with repeated coverage on descending or ascending orbits or where the terrain is flat and the targets nonoriented, coverage may be sufficient to provide data for change detection. Examples of Los Angeles and the Everglades of Florida help develop and support the argument.
The elevation gradients affecting tropical forest stand characteristics are presently studied in light of multipolarization airborne SAR data. A 'rubber sheeting' computer code was used to georeference the SAR data sets to the digital elevation data. The TOPO code from NASA's NSTL generated the terrain slope and aspect angle data from the terrain elevation data set; computed local incidence angles were used to delete those data areas that were shadowed, and to produce local incidence angle data that were not shadowed. The results obtained demonstrate that the SAR data are related to the elevation gradient.
Synthetic data is a powerful tool to generate large amounts of training data for machine learning models. The methods outlined in this report will be used to retrain the deep learning classifier for increased accuracy. Synthetic data will be useful to address the natural class imbalance between the different categories in the original ML work. Additionally, these tools will be applied for a variety of signal analysis methods that would use signals with a known signal-to-noise ratio for validation and testing.
A new GAL4-based feed-forward loop circuit enhances β-glucuronidase (GUS) reporter gene expression in leaves and stems of stably transformed sugarcane plants.
Synthetic aperture radar (SAR) instruments on spacecraft are capable of producing huge quantities of data. Onboard lossy data compression is commonly used to reduce the burden on the communication link. In this paper an overview is given of various SAR data compression techniques, along with an assessment of how much improvement is possible (and practical) and how to approach the problem of obtaining it. Synthetic aperture radar (SAR) instruments on spacecraft are capable of acquiring huge quantities of data. As a result, the available downlink rate and onboard storage capacity can be limiting factors in mission design for spacecraft with SAR instruments. This is true both for Earth-orbiting missions and missions to more distant targets such as Venus, Titan, and Europa. (Of course for missions beyond Earth orbit downlink rates are much lower and thus potentially much more limiting.) Typically spacecraft with SAR instruments use some form of data compression in order to reduce the storage size and/or downlink rate necessary to accommodate the SAR data. Our aim here is to give an overview of SAR data compression strategies that have been considered, and to assess the prospects for additional improvements.
Synthetic aperture radar data was used to construct an estimation algorithm for development of information on long waves. The evolution of chaotic dynamic systems was also explored.
Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.
Broad absorption line quasars (BALs) exhibit blueshifted absorption relative to a number of their prominent broad emission features. These absorption features can contribute to quasar redshift errors and add absorption to the Lyman-α (Lyα) forest that is unrelated to large-scale structure. We present a detailed analysis of the impact of BALs on the Baryon Acoustic Oscillation (BAO) results with the Lyα forest from the first year of data from the Dark Energy Spectroscopic Instrument (DESI). The baseline strategy for the first year analysis is to mask all pixels associated with all BAL absorption features that fall within the wavelength region used to measure the forest. We explore a range of alternate masking strategies and demonstrate that these changes have minimal impact on the BAO measurements with both DESI data and synthetic data. This includes when we mask the BAL features associated with emission lines outside of the forest region to minimize their contribution to redshift errors. We identify differences in the properties of BALs in the synthetic datasets relative to the observational data, as well as use the synthetic observations to characterize the completeness of the BAL identification algorithm, and demonstrate that incompleteness and differences in the BALs between real and synthetic data also do not impact the BAO results for the Lyα forest.
Acoustic Resonance Spectroscopy (ARS) is highly sensitive to structural properties such as material, geometry, and environmental conditions; as a consequence, it can noninvasively measure internal properties that are unobservable by most other methods. Because of its sensing capabilities and low implementation cost and complexity, ARS has potential as a paradigm shift in noninvasive sensing, characterization, and monitoring applications. However, extracting specific properties from ARS measurements, comprising the vibration spectrum of a test object, is challenging due to the sensitivity of the spectra to other structural changes not being measured, e.g. manufacturing tolerances, component coupling, environmental variation, etc. Neural Networks are promising tools for identifying trends in ARS measurements, but their training typically requires large datasets, which are often impractical to obtain for real-world systems. Synthetic data can be simulated efficiently, but discrepancies between synthetic and real-world data frequently lead to poor generalization when testing on the real-world data. We propose a novel ARS model training framework that enables networks trained exclusively on synthetic ARS data to generalize effectively to real-world measurements. Our approach leverages the Correlation Alignment (CORAL) technique to enforce the extraction of features common to both synthetic and real-world domains. As a case study, we demonstrate noninvasive ARS-based pressure measurements in sealed systems. Finite element method (FEM) simulations were used to generate synthetic training data across diverse vessel configurations and pressure conditions, and model performance was then tested on real-world measurements. We demonstrate that robust machine learning models for ARS can be developed without large real-world datasets, significantly broadening the applicability of ARS for noninvasive sensing. Moreover, the approach is extensible to other sensing modalities where synthetic data are abundant but real-world data are limited.