Search NASASearch

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Generalizing synthetic data-trained acoustic predictive models to real-world measurements

Acoustic Resonance Spectroscopy (ARS) is highly sensitive to structural properties such as material, geometry, and environmental conditions; as a consequence, it can noninvasively measure internal properties that are unobservable by most other methods. Because of its sensing capabilities and low implementation cost and complexity, ARS has potential as a paradigm shift in noninvasive sensing, characterization, and monitoring applications. However, extracting specific properties from ARS measurements, comprising the vibration spectrum of a test object, is challenging due to the sensitivity of the spectra to other structural changes not being measured, e.g. manufacturing tolerances, component coupling, environmental variation, etc. Neural Networks are promising tools for identifying trends in ARS measurements, but their training typically requires large datasets, which are often impractical to obtain for real-world systems. Synthetic data can be simulated efficiently, but discrepancies between synthetic and real-world data frequently lead to poor generalization when testing on the real-world data. We propose a novel ARS model training framework that enables networks trained exclusively on synthetic ARS data to generalize effectively to real-world measurements. Our approach leverages the Correlation Alignment (CORAL) technique to enforce the extraction of features common to both synthetic and real-world domains. As a case study, we demonstrate noninvasive ARS-based pressure measurements in sealed systems. Finite element method (FEM) simulations were used to generate synthetic training data across diverse vessel configurations and pressure conditions, and model performance was then tested on real-world measurements. We demonstrate that robust machine learning models for ARS can be developed without large real-world datasets, significantly broadening the applicability of ARS for noninvasive sensing. Moreover, the approach is extensible to other sensing modalities where synthetic data are abundant but real-world data are limited.

36 MATERIALS SCIENCE

Improved Earthquake Source Parameters with 3D Wavespeed Models in California and Nevada

Seismic tomography harnesses earthquake data to explore the inaccessible structure of the Earth. Adjoint waveform tomography (AWT), a method of seismic tomography, updates the tomographic model by optimizing the fit between observed earthquake data and synthetic waveforms. The synthetic data are calculated by solving the wave equation through a given 3D model. An important requirement to calculating synthetics is the source information (location, centroid time, depth, and moment tensor). Errors in source information affect the quality of the synthetics produced, which in turn can limit how structure can be inferred in the AWT workflow. Here, to test the effect of updating source information, we used MTTime (Chiang, 2020), a time-domain full-waveform moment tensor inversion code, to calculate the moment tensors and depths of 118 earthquakes that occurred in California and Nevada over a 20-yr period. We calculated 3D Green’s functions using a 3D seismic wavespeed model of California and Nevada (Doody et al., 2023b). We show that the inverted solutions provide better waveform fits than the Global Centroid Moment Tensor catalog and increase usable, well-correlated data by up to 7%. Therefore, we argue that recalculating source parameters should be considered in AWT workflows, particularly for smaller magnitude events (⁠M w > 5.0).

58 GEOSCIENCES

An analysis of the focusing performance of Magellan SAR image data

Synthetic aperture radar (SAR) focusing can be achieved either based on accurate ephemeris data or on an autofocusing process. For the Magellan project, such a decision must be made in the early phase of Magellan SAR system design. The analysis of the emphemeris requirement is complicated. The analysis given by the author leads to the conclusion that empheris data obtained from the Magellan navigation system provide sufficient accuracy to meet the Magellan image resolution requirement.

Jin, Michael Y.

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]

Principle Component Analysis of AIRS and CrIS Data

Synthetic Eigen Vectors (EV) used for the statistical analysis of the PC reconstruction residual of large ensembles of data are a novel tool for the analysis of data from hyperspectral infrared sounders like the Atmospheric Infrared Sounder (AIRS) on the EOS Aqua and the Cross-track Infrared Sounder (CrIS) on the SUOMI polar orbiting satellites. Unlike empirical EV, which are derived from the observed spectra, the synthetic EV are derived from a large ensemble of spectra which are calculated assuming that, given a state of the atmosphere, the spectra created by the instrument can be accurately calculated. The synthetic EV are then used to reconstruct the observed spectra. The analysis of the differences between the observed spectra and the reconstructed spectra for Simultaneous Nadir Overpasses of tropical oceans reveals unexpected differences at the more than 200 mK level under relatively clear conditions, particularly in the mid-wave water vapor channels of CrIS. The repeatability of these differences using independently trained SEV and results from different years appears to rule out inconsistencies in the radiative transfer algorithm or the data simulation. The reasons for these discrepancies are under evaluation.

infrared

Radiation effects response data for synthetic organic insulation and dielectrics

Existing radiation response data for 130 materials of the synthetic organic insulation and dielectric class are analyzed, and thresholds and 25-percent change dose levels for these materials are presented. Both the lowest reported threshold dose (LTD) and the 25-percent change dose level are found to vary widely among the different insulators (used as spacecraft components). The LTD is tabulated to indicate the levels where radiation effects become apparent. These level are of use in setting a lower limit where changes can be expected, and could be used to exempt synthetic insulation and dielectrics from radiation considerations when the environments are substantially below this limit. Cautions to be observed in applying the data to current problems are presented.

Bouquet, F. L.

Data efficiency assessment of generative adversarial networks in energy applications

This study investigates the data requirements of generative artificial intelligence (AI), particularly generative adversarial networks (GANs), for reliable data augmentation in energy applications. Generative AI, though seen as a solution to data limitations, requires substantial data to learn meaningful distributions—a challenge often overlooked. This study addresses the challenge through synthetic data generation for critical heat flux (CHF) and power grid demand, focusing on renewable and nuclear energy. Two variants of GAN employed are conditional GAN (cGAN) and Wasserstein GAN (wGAN). Our findings include the strong dependency of GAN on data size, with performance declining on smaller datasets and varying performance when generalizing to unseen experiments. Mass flux and heated length significantly influence CHF predictions. wGAN is more robust to feature exclusion, making it suitable for constrained synthetic data generation. In energy demand forecasting, wGAN performed well for solar, wind, and load predictions. Longer lookback hours and larger datasets improved predictions, especially for load power. Seasonal variations posed challenges, with wGAN achieving a relatively high error of Root Mean Squared Error (RMSE) of 0.32 for load power prediction, compared to RMSE of 0.07 under same-season conditions. Feature exclusions impacted cGAN the most, while wGAN showed greater robustness. This study concludes that, while generative AI is effective for data augmentation, it requires substantial data and careful training to generate realistic synthetic data and generalize to new experiments in engineering applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

SSTDR and FDR Detection of Un-Energized and Energized Cable Anomalies Including Thermal Degradation Using Machine Learning

Historically, cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, continued use of these cables must shift to a performance-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. A variety of cable tests are available and are commonly applied during outages when the cables can be taken out of service. Frequency domain reflectometry (FDR) is one of these test methods that is being more broadly accepted and used because it not only detects anomalies along the cable with a low-voltage signal that does not stress the cable insulation, but the technique also locates the anomalies. This supports follow-up local inspection and local repair or partial replacement of a damaged cable segment. Currently, FDR testing is only applied to cables that are taken out of service since the test instrument would be damaged by operational voltages. A related technology that has found some acceptance in the aircraft and rail industry is spread spectrum time domain reflectometry (SSTDR). This technology has been implemented with a custom commercial instrument by LiveWire Innovation that is designed to operate on live cables up to 1000 volts and with a bandwidth of 48 MHz. Initial evaluation by the Pacific Northwest National Laboratory (PNNL) of the Live Wire system indicated that a broader bandwidth (BW) SSTDR may be better for many kinds of flaws. This led PNNL to develop an SSTDR laboratory instrument suitable for tests up to 500 MHz bandwidth. Testing on energized cables is also desirable for online monitoring systems so an inductive clamshell coupler was developed that allows energized cables to be tested up to at least 5 kV and likely higher voltage levels. Dielectric spectroscopy and tan delta testing plus various laboratory destructive tests were included in this data acquisition campaign directed to feed a machine learning (ML) study. With these kinds of developments, online energized cable tests may be possible with industrial adoption of such hardware advances but it will be completely impractical to have highly skilled data analysts continually examine these complex signals for indications of damage or compromised conditions. If online testing is to be implemented in new test hardware, it must be accompanied by software that can interpret the signals and alert plant operators of changing or degraded conditions. The thermally aged, shielded cable investigated here was separately treated for ML analysis. Visual analysis of electrical data showed generally increasing peaks where the cable entered and exited the oven. These peaks were not exactly aligned with expected locations, but these differences were attributed to velocity of propagation calibration errors. Only supervised ML was applied to the thermally aged data as this data was only available shortly before the committed publication date of this report. The supervised ML was structured to divide the 0 to 70-day responses as ‘normal’ from 0 to 35 days or ‘anomalous’ from 36 to 70 days, based on cable tensile elongation at break (EAB) insulation characterization. Using 80% of the data for training and 20% for testing, the supervised ML predicted normal versus anomalous was 70% accurate. Important conclusions include: • Accuracy to predict the presence of cable damage is improved from the 2023 effort by more training data. Weighted accuracies for comparisons among the instruments ranged from 67 to 89 % for unsupervised ML and 71 to 99% for supervised ML. • Based on the synthetic data tests, the unsupervised models are more generalizable to unseen anomalies. The Multi-Layer Perceptron classifier (MLP) model reported as high as 99.7% accuracy on the test data, but this dropped to 58.3% when tested on the synthetic data. In contrast, the unsupervised Pointwise model only achieved 89.7% accuracy on the experimental data but reported 78.3% accuracy on the synthetic data. • The best anomaly indicators are higher frequency (400 MHz BW) FDR data. Other tests may be interesting but for this study, this was the best predicter.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.

Hybrid Flush and Synthetic Air Data Filter for Entry Vehicle Atmospheric State Estimation

A hybrid flush/synthetic air data sensing filter utilizing Kalman-Schmidt and Rach-Tung-Striebel smoothers is developed to obtain entry vehicle atmosphere estimates. The filter/smoother blends information from pressure sensors distributed on the heatshield with measurements of the vehicle aerodynamic forces and moments computed from mass properties and inertial measurement unit data, and prior estimates of the atmosphere. The filter produces estimates of the atmospheric conditions along the entry trajectory, and systematic error estimates to reconcile differences between the pressure and aerodynamic data sources. The filter is applied to data acquired during the Mars Science Laboratory and Mars 2020 entry, descent, and landing at Gale crater and at Jezero crater, respectively. The results show that the hybrid filter produces estimates of the freestream flight condition with lower uncertainty than either the flush or synthetic air data algorithms. The filter accomplishes this result by incorporating additional data and computing estimates of systematic error parameters in the pressure data and the aerodynamic model to further reduce the uncertainties.

Christopher D. Karlgaard

Calibration verification for stochastic agent-based disease spread models

Accurate disease spread modeling is crucial for identifying the severity of outbreaks and planning effective mitigation efforts. To be reliable when applied to new outbreaks, model calibration techniques must be robust. However, current methods frequently forgo calibration verification (a stand-alone process evaluating the calibration procedure) and instead use overall model validation (a process comparing calibrated model results to data) to check calibration processes, which may conceal errors in calibration. In this work, we develop a stochastic agent-based disease spread model to act as a testing environment as we test two calibration methods using simulation-based calibration, which is a synthetic data calibration verification method. The first calibration method is a Bayesian inference approach using an empirically-constructed likelihood and Markov chain Monte Carlo (MCMC) sampling, while the second method is a likelihood-free approach using approximate Bayesian computation (ABC). Simulation-based calibration suggests that there are challenges with the empirical likelihood calculation used in the first calibration method in this context. These issues are alleviated in the ABC approach. Despite these challenges, we note that the first calibration method performs well in a synthetic data model validation test similar to those common in disease spread modeling literature. We conclude that stand-alone calibration verification using synthetic data may benefit epidemiological researchers in identifying model calibration challenges that may be difficult to identify with other commonly used model validation techniques.

60 APPLIED LIFE SCIENCES

Fan Beam Emission Tomography for Estimating Scalar Properties in Laminar Flames

A new method of estimating temperatures and gas species concentrations (CO2 and H2O) in a laminar flame is reported. The path-integrated, spectral radiation intensities emitted from a laminar flame at multiple wavelengths and view angles are calculated using a narrow band radiation model. Synthetic data, in the form of radial profiles of temperature and gas concentrations, are used in these calculations. The calculations mimic measurements that would theoretically be obtained using a mid-infrared spectrometer with a scanner. The path integrated spectral radiation intensities are deconvoluted using a maximum likelihood estimation method in conjunction with an iterative scheme. The deconvolution algorithm accounts for the self-absorption of radiation by the intervening gases, and provides the local temperature and gas species concentrations. The deconvoluted temperatures and gas concentrations are compared with the synthetic data used for calculating the spectral radiation intensities. The deconvoluted temperatures and gas species concentrations are within 0.5 % of the synthetic data. The deconvolution algorithm is expected to provide combustion researchers with an easy method of obtaining the radial profiles of major gas species concentrations and temperatures in laminar flames non-intrusively using a mid-infrared spectrometer with a scanner.

Lim, Jongmook

The two-way time synchronization system via a satellite voice channel

A newly developed two-way time synchronization system is described in this paper. The system uses one voice channel at a SCPC satellite digital communication earth station, whose bandwidth is only 45 kHz, thus saving satellite resources greatly. The system is composed of one master station and one or several, up to sixty-two, secondary stations. The master and secondary stations are equipped with the same equipment, including a set of timing equipment, a synthetic data terminal for time synchronizing, and a interface unit between the data terminal and the satellite earth station. The synthetic data terminal for time synchronization also has an IRIG-B code generator and a translator. The data terminal of master station is the key part of whole system. The system synchronization process is full automatic, which is controlled by the master station. Employing an autoscanning technique and conversational mode, the system accomplishes the following tasks: linking up liaison with each secondary station in turn, establishing a coarse time synchronization, calibrating date (years, months, days) and time of day (hours, minutes, seconds), precisely measuring the time difference between local station and the opposite station, exchanging measurement data, statistically processing the data, rejecting error terms, printing the data, calculating the clock difference and correcting the phase, thus realizing real-time synchronization from one point to multiple points. We also designed an adaptive phase circuit to eliminate the phase ambiguity of the PSK demodulator. The experiments have shown that the time synchronization accuracy is better than 2 mu S. The system has been put into regular operation.

Heng-Qiu, Zheng

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis

Synthetic Hyperspectral Data for Global Water Quality Algorithm Development

Eutrophication and increasing prevalence of potentially toxic algal blooms (cyanoHABs) among global inland water bodies have become a major ecological concern and require direct attention. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local aquatic processes, to spatially resolved global products. Planned aquatic biogeochemistry remote sensing data products from hyperspectral imagers such as NASA’s Surface Biology and Geology (SBG) mission and relevant aquatic sensor sensitivity precursor airborne imaging spectrometer data provide unprecedented radiometric resolution and sensor sensitivity for characterizing complex aquatic ecosystems. However, scarcity of high-quality freshwater in-situ optical data hinders our capability to develop and validate robust retrieval algorithms. A state-of-the-art synthetic dataset of paired top-of-atmosphere, bottom-of-atmosphere, and optical and biogeophysical data was developed through radiative transfer modeling to simulate natural freshwater ecosystems. A synthetic or precursor dataset for SBG is being used to train robust machine learning models to derive water quality products pertinent to SBG mission objectives. The dataset is also used to show the potential of performing vigorous aquatic sensitivity studies and explored pathways for how best to optimize hyperspectral data for machine learning development. A processing pipeline and resultant global synthetic/precursor dataset for inland waters is presented to establish the innovation for water quality studies of inland waters globally. Optical Society of America Imaging and Applied Optics Congress, Hyperspectral Imaging and Sounding of the Environment (OSA HISE) Meeting, 19-23 July 2021, Virtual Meeting, https://www.osa.org/enus/meetings/osa_meetings/optical_sensors_and_sensing_congress/program/hyperspectral_imaging_and_sounding_of_the_environm/

Synthetic