Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Hierarchical Bayesian Modeling for Cosmology: Can NPE reliably replace MCMC?

Hierarchical neural posterior estimation has its place Hierarchical Bayesian Modeling (HBM) combined with MCMC algorithms has been shown to provide more robust and accurate inference for real-world phenomena in which nature takes a nested form. However, MCMC-based inference can be computationally expensive, and its performance often suffers for complex posterior geometries. These costs are especially pertinent for HBM. Studies have recently demonstrated the potential for a flexible, expressive, and amortized hierarchical neural posterior estimator (HNPE) built on Normalizing Flows. These studies have mostly been performed on simple datasets, or they focus on a single parameter from each level of the hierarchy. A systematic study analyzing how both hierarchical methods compare for more complex and realistic datasets is necessary before applying HNPE for scientific measurements. Here, we re-explore the theory behind HNPE and conduct comparative numerical experiments of HNPE and MCMC-based HBM methods on real and synthetic data, including strong gravitational lensing simulations. In particular, we use a suite of diagnostics to show trade-offs in terms of accuracy, precision, time to train or sample, reproducibility, and the need for expert domain knowledge. Especially for higher dimensional and complex posteriors, HNPE is expected to drastically improve on time for inference, accuracy, and precision with an upfront training time cost.

Hur, Rachel [Chicago U.] (ORCID:000900089890445X)↗

Red Noise–based False Alarm Thresholds for Astrophysical Periodograms via Whittle’s Approximation to the Likelihood

Astronomers who search for periodic signals using Lomb–Scargle periodograms rely on false alarm level (FAL) estimates to identify statistically significant peaks. Although FALs are often calculated from white noise models, many astronomical time series suffer from red noise. Prewhitening is a statistical technique in which a continuum model is subtracted from the log power spectrum estimate, after which the observer can proceed with a white-noise treatment. Here we present a prewhitening-based method of calculating frequency-dependent FALs. We fit power laws and autoregressive models of order 1 to each Lomb–Scargle periodogram by minimizing the Whittle approximation to the negative log-likelihood (NLL), then calculate FALs based on the best-fit model power spectrum. Our technique is a novel extension of the Whittle NLL to datasets with uneven time sampling. We demonstrate FAL calculations using observations of α Cen B, GJ 581, HD 192310, synthetic data from the radial velocity (RV) fitting challenge, and Kepler observations of a differential rotator. The Kepler data analysis shows that only true rotation signals are detected by red noise FALs, while white noise FALs suggest all spurious peaks in the low-frequency range are significant. A high-frequency sinusoid injected into α Cen B logR$'$ HK observations exceeds the 1% red noise FAL despite having only 8.9% of the power of the dominant rotation signal. In a periodogram of HD 192310 RVs, peaks associated with differential rotation and planets are detected against the 5% red noise FAL without iterative model fitting or subtraction. The software for calculating red noise–based FALs is available on GitHub.

Astrostatistics (1882)↗

Cosmic Shear Analysis of the DECam Local Volume Exploration Survey

We forecast cosmological constraints and develop a cosmic shear analysis pipeline for the DECam Local Volume Exploration Survey (DELVE). We test the effects of two different intrinsic alignment frameworks (TATT and NLA) on synthetic data vectors. In addition, we examine the impact of baryon contamination and determine the necessary scale cuts to reduce its influence. We find the forecast results to be as constraining as the DES Y3 cosmological parameter measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Convolution Neural Network for Voltage Event Classification at a Photovoltaic Inverter

This paper presents a convolutional neural network (CNN) developed to identify voltage events in photovoltaic (PV) inverters. The CNN is trained on synthetic data generated using the IEEE 13-bus distribution feeder model and evaluated on field measured data collected from Energy Northwest’s Horn Rapids Solar, Storage, and Training (HRSST) facility. The study focuses on two common voltage events: faults and voltage sags. The CNN is configured to analyze voltage and current waveforms from three-phase PV systems, demonstrating excellent accuracy during training. Field data from the HRSST facility is employed to assess its real-world performance, where the CNN achieves perfect identification of faults and voltage sags in a sample of nine events. This work highlights the potential of the proposed method to enhance PV protection schemes, providing a robust foundation for improved voltage event detection and grid reliability.

Cornachione, Matthew A.↗

A Visual Analytic Platform for Interactive Validation of Human Mobility Simulations

Human mobility insights guide domain experts in an array of decisions, including critical infrastructure design, disaster response, epidemic modeling, national security, and policy making. Due to the inherent noise and privacy concerns in real-world individual-level mobility data, it is often preferred to leverage simulators that generate synthetic mobility data instead. However, it is critical to inspect and validate the output of such simulators to ensure the synthetic data is aligned with the characteristics of the population and the area of interest known to domain experts. While there exist many quantitative approaches for validating synthetic data, we argue it is also important to also validate such data qualitatively to capture aspects that are known to domain experts but difficult to quantify. In this work, we demonstrate a visual analytic platform that empowers domain experts to interact with their simulation outputs along spatial and temporal dimensions. By augmenting automated techniques and human skills, our visual analytic platform is a step towards interactive capabilities for model steering and quality control of mobility simulators.

Monadjemi, Shayan↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

An ultra-fast method for generating synthetic down-scattered neutron data for inertial confinement fusion implosions

In inertial confinement fusion experiments at the National Ignition Facility, asymmetries are probed by a variety of neutron diagnostics, including neutron imaging systems, real-time neutron activation diagnostics (RTNADs), and neutron spectrometers. It is often useful to generate synthetic data based on these diagnostics to validate and tune models. However, current methods of doing so using Monte Carlo particle tracing are time-consuming. In this paper, an ultra-fast method is presented for generating synthetic neutron images, RTNAD data, and spectrometry data using line integrals and 3D convolutions. While it does not contain as much physics as particle tracing codes, it is thousands of times faster and produces nearly identical data. This enables analysis techniques that depend on generating large amounts of synthetic data, which will prove very useful for the study of asymmetries going forward.

Deuterium↗

Linking Threat Agents to Targeted Organizations: A Pipeline for Enhanced Cybersecurity Risk Metrics

In this study, we present a methodology leveraging Large Language Models (LLMs) to transform Cybersecurity Threat Intelligence (CTI) narratives into actionable insights for individual organizations. Our approach automates the extraction of machine-readable adversary SKRAM (Skills, Knowledge, Resources, Authorities, and Motivation) attributes from open-source reports, extending LLM utility beyond typical interactions. This innovation enables precise, automated assessments of cybersecurity risks posed by various adversaries. Using a chain-of-thought and multi-shot prompting strategy, our methodology advances the automation of cybersecurity feature extraction for new machine-learning models that predict the risk of adversary targeting. This approach is refined using a substantial dataset of over 150 analyst-validated threat reports and synthetic organizational data from 900 companies. Here, by bootstrapping the training data with a rule-based heuristic over synthetic data, we have developed a high-accuracy machine-learning model that allows entities to dynamically prioritize threats and defensive actions.

Cyber Threat Intelligence↗

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins↗

A Data-Driven Method for Synthetic Extreme Weather Generation and Solar Impact Assessment: Preprint

High-resolution, high-fidelity weather datasets are essential for testing and evaluating the resilience of power systems, particularly under extreme weather conditions. However, existing extreme weather datasets are typically derived from historical events that are localized and may lack the spatial and temporal resolution or scenario diversity needed to test largescale power systems. In this work, we propose a synthetic extreme weather simulation approach capable of generating targeted extreme events, such as hurricanes, using publicly available data sources. Preliminary results demonstrate the impact of a simulated Category 1 hurricane on renewable generation and critical infrastructure in California. The work aims to provide a flexible approach for creating multiple types of extreme weather scenarios across different regions, enabling comprehensive system stress testing, training, and resilience assessment.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Orbital-Radar v1.0.0: a tool to transform suborbital radar observations to synthetic EarthCARE cloud radar data

The Earth Cloud, Aerosol and Radiation Explorer (EarthCARE) satellite developed by the European Space Agency (ESA) and the Japan Aerospace Exploration Agency (JAXA) launched in May 2024 carries a novel 94 GHz cloud profiling radar (CPR) with Doppler capability. This work describes the open-source instrument simulator Orbital-Radar, which transforms high-resolution radar data from field observations or forward simulations of numerical models to CPR primary measurements and uncertainties. The transformation accounts for sampling geometry and surface effects. We demonstrate Orbital-Radar's ability to provide realistic CPR views of typical cloud and precipitation scenes. The presented case studies show small-scale convection, marine stratus clouds, and Arctic mixed-phase cloud cases. These results provide valuable insights into the capabilities and challenges of the EarthCARE CPR mission and its advantages over the CloudSat CPR. Finally, Orbital-Radar allows for evaluating kilometre-scale numerical weather prediction models with EarthCARE CPR observations. So, Orbital-Radar can generate calibration and validation (Cal/Val) data sets already pre-launch. Nevertheless, an evaluation of synthetic CPR output data to accurate EarthCARE CPR data is missing.

54 ENVIRONMENTAL SCIENCES↗

Data‐Efficient Generation of Synthetic Microstructures of Polymer‐Bonded Energetic Material With Fine‐Tuned Stable Diffusion

Among current deep learning approaches for synthetic image generation, diffusion-based models stand out in terms of algorithmic stability and ability to retain high-fidelity image features with detailed resolution. Here, in this work, we employ Dreambooth, a method for fine-tuning Stable Diffusion, on X-ray CT images of microstructure of the polymer-bonded form (PBX) of a commonly used high explosive, Pentaerythritol tetranitrate (PETN), which yields generative models for creating synthetic PBX images. The models developed here represent five classes (or ‘lots’) of microstructures and demonstrate successful generation of images of each class with high fidelity, as verified by computed classification accuracy of ∼ 94% or higher. Data augmentation afforded by such image synthesis can be used to more reliably decipher underlying statistics, build processing-structure correlations, recognize off-normal structural anomalies, and identify age-related changes. Ideas related to converting image data into appropriate density mapping and performing mesoscale simulation or surrogate modeling of detonation are also discussed.

Dreambooth↗

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES↗

Machine learning approaches for crystallographic classification from synthetic 2D X-ray diffraction data

Crystallographic structure identification is crucial for understanding material properties; however, current methodologies often depend on labor-intensive and time-consuming analyses of 2D X-ray diffraction (XRD) patterns. To address these limitations, this study employs synthetic 2D XRD patterns combined with deep learning (DL) techniques to enable automated and high-throughput classification of the seven crystal systems and 230 space groups. We introduce the novel Auto Diffraction Pipeline, designed to generate synthetic 2D XRD spot patterns from crystallographic information files under diverse conditions, including varying zone axes, atomic substitution, atomic depletion and mechanical loading. These conditions enhance the realism of synthetic data, mitigating the scarcity of experimental datasets and enabling the creation of large representative training sets. Convolutional neural networks were trained and validated on these synthetic datasets to classify crystallographic structures across multiple scenarios. Our results demonstrate that integrating synthetic 2D XRD patterns with DL facilitates rapid, accurate and automated crystallographic classification, promoting the wider adoption of data-driven approaches in materials science.

Shahnazari, Ayoub [Univ. of Rochester, NY (United ↗

Validation of the DESI 2024 Lyα forest BAO analysis using synthetic datasets

The first year of data from the Dark Energy Spectroscopic Instrument (DESI) contains the largest set of Lyman-α (Lyα) forest spectra ever observed. This data, collected in the DESI Data Release 1 (DR1) sample, has been used to measure the Baryon Acoustic Oscillation (BAO) feature at redshift z = 2.33. In this work, we use a set of 150 synthetic realizations of DESI DR1 to validate the DESI 2024 Lyα forest BAO measurement presented in [1]. The synthetic data sets are based on Gaussian random fields using the log-normal approximation. We produce realistic synthetic DESI spectra that include all major contaminants affecting the Lyα forest. The synthetic data sets span a redshift range 1.8 < z < 3.8, and are analyzed using the same framework and pipeline used for the DESI 2024 Lyα forest BAO measurement. To measure BAO, we use both the Lyα auto-correlation and its cross-correlation with quasar positions. We use the mean of correlation functions from the set of DESI DR1 realizations to show that our model is able to recover unbiased measurements of the BAO position. We also fit each mock individually and study the population of BAO fits in order to validate BAO uncertainties and test our method for estimating the covariance matrix of the Lyα forest correlation functions. Finally, we discuss the implications of our results and identify the needs for the next generation of Lyα forest synthetic data sets, with the top priority being to simulate the effect of BAO broadening due to non-linear evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

Monte Carlo toolkit for designing and validating step-range-filter spectrometer designs

Here, we present a Monte Carlo toolkit for validating step range filter (SRF) spectrometer designs. Geant4 is used to transport charged particles through the SRF filters to generate synthetic SRF data that include realistic CR-39 effects. Synthetic SRF spectra generated by this method inherently account for instrument response and allow for the quantification of SRF performance before shots. The usefulness of this toolkit is demonstrated through its application to a number of problems. A new broadband SRF for the ∼10 MeV wide 3He3He proton spectrum is validated, and an analysis method for analyzing 3He3He-p SRF data that accounts for instrument response is put forth. In addition, an SRF design for the compact recoil-proton spectrometer (CRS) on the Z-machine is validated. Finally, a new calibration technique for the DD-p SRF is proposed and validated.

Johnson, T. M. (ORCID:0000000193032949)↗