Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Biopolymer-Templated Titania Film Formation for Nanostructured Coatings Revealed by Machine Learning-Supported Time-Resolved Analysis

This study presents a machine learning approach to derive the film formation of biopolymer-templated titania nanostructures during spray deposition, in combination with in situ grazing-incidence small-angle X-ray scattering (GISAXS). A neural network trained on synthetic GISAXS data directly predicts domain-size distributions from experimental two-dimensional scattering patterns, capturing the full kinetics of nanostructure evolution with high temporal resolution. The predictions reveal hierarchical size distributions and periodic growth features, consistent with layer-by-layer spray deposition and validated by complementary scanning electron microscopy (SEM) imaging. Quantitative comparison with conventional parametric GISAXS fits shows good qualitative agreement, with systematic differences explained by domain-shape assumptions and resolved by applying a geometric scaling factor. Simulated SEM-like surfaces derived from neural network outputs reproduce the porous, foam-like nanoscale morphology observed experimentally, reinforcing the method’s credibility. This integrated approach enables real-time, nondestructive, statistically averaged monitoring of bulk nanostructure development in functional coatings, offering a scalable methodology to accelerate the characterization and process control of sustainably manufactured nanostructured titania films for energy-related applications such as photocatalysis and photovoltaics.

Heger, JulianEliah↗

Machine learning inversion of interatomic force constants from single-crystal inelastic neutron scattering

Atomic vibrations govern many macroscopic properties of materials, but experiments to comprehensively probe them remain challenging. Inelastic neutron scattering (INS) is a powerful technique to map phonon dispersions in crystals, especially when leveraging modern time-of-flight (ToF) spectrometers with large detectors. However, efficiently and robustly extracting interatomic force constants (FCs) parameterizing phonon dynamics from experimental spectra remains a bottleneck due to the complexity and high dimensionality of ToF INS datasets. Here, we present a machine learning approach for the direct inversion of FCs from single-crystal INS measurements. The framework leverages synthetic training data generated using universal machine-learned force fields and an efficient physics-based forward model. We benchmark two neural architectures–one emphasizing structured latent representation learning and the other direct, supervised spectral regression–across simulated datasets for two materials under idealized and noisy conditions. The latent-representation model is subsequently applied to experimental single-crystal INS data on germanium. The model is shown to reproduce FCs derived from both first-principles simulations and from iterative optimization, and furthermore achieves reliable inference even from sparse, single-orientation measurements representing short data acquisitions. Analysis of the learned latent space reveals semantically continuous and physically interpretable encodings that support strong cross-domain generalization. By bridging theoretical and experimental domains, we establish a path toward rapid inversion of experimental spectra and data-driven interpretation of temperature-dependent lattice dynamics.

42 ENGINEERING↗

Distortions in charged-particle images of laser direct-drive inertial confinement fusion implosions

Energetic charged particles generated by inertial confinement fusion (ICF) implosions encode information about the spatial morphology of the hotspot and dense fuel during the time of peak fusion reactions. The knock-on deuteron imager (KoDI) was developed at the Omega Laser Facility to image these particles in order to diagnose low-mode asymmetries in the hotspot and dense fuel layer of cryogenic deuterium–tritium ICF implosions. However, the images collected are distorted in several ways that prevent reconstruction of the deuteron source. In this paper, we describe these distortions and a series of attempts to mitigate or compensate for them. We present several potential mechanisms for the distortions, including a new model for scattering of charged particles in filamentary electric or magnetic fields surrounding the implosion. Particle-tracing is used to create synthetic KoDI data based on the filamentary field model that reproduces the main experimentally observed image distortions. We conclude that the filamentary scattering model best matches the observed image distortions. Finally, we discuss potential impacts of filamentary fields on other charged-particle diagnostics.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Evaluation of a generalized least squares algorithm for infrasound beamforming with coherent background noise

Infrasonic signals of interest can occur during periods with persistent, coherent, background noise, which may be natural or anthropogenic. For high signal-to-noise (SNR) ratio transient signals, an ‘overprinting’ of the coherent background may occur, and the signal may still be detected. However, this approach fails for low SNR signals of interest, which may be obscured by coherent noise. An infrasound beamforming method based on generalized least squares (GLS) is investigated for detecting transient signals of interest in the presence of coherent and incoherent background noise. This approach relies on an estimate of the noise covariance, captured in a covariance matrix, to effectively null contributions to the array response from noisy directions of arrival. Synthetic array data is used to investigate the performance of the GLS beamformer compared to the Bartlett beamformer when coherent and incoherent backgrounds are present. Additionally, the effects of array element number and relative strength of the interfering signal on the GLS estimates is investigated. GLS empirical area under the curve estimates suggest that the beamformer can recover coherent power for a signal of interest lower in amplitude than the coherent background, but this effectiveness degrades more quickly with SNR for a four element array compared to a six or eight element infrasound array. Finally, infrasound from the Forensic Surface Experiment, a bolide signal observed at IMS array I37NO, and a volcanic signal recorded at the Alaska Volcano Observatory array ADKI are used to evaluate GLS performance on recorded data. A ten minute window was used to capture the background noise, and the coherent background signal was nulled in all three examples.

58 GEOSCIENCES↗

NNFDivergence

The code implements f divergence regularization for neural networks in the Python-based Pytorch framework. The methods are the main focus but the repository will also contain examples that operate on purely synthetic "toy" data or on openly available, public data from NASA.

Klein, Natalie [@lanl]↗

Assessing the Value of Seismic Amplitude Versus Offset (AVO) Attributes for CO2 Storage Project Using a Bayesian Network Model for Decision Support

Attributes versus offset (AVO) are a set of measurements to analyze how the characteristics of reflected seismic waves change as a function of the offset. It can be useful for monitoring CO2 storage sites because the presence of leaked CO2 into the overlying aquifer can change the properties of the rocks and pore fluids composition that can alter the way seismic waves reflect and their amplitudes. The time-lapse changes in AVO attributes derived from repeat seismic surveys can help identify anomalies or shifts in the subsurface that could potentially be used as an indicator for CO2 leak detection. Our study leverages multiple seismic attributes derived from synthetic seismic data and Bayesian network model to quantify the probability of leak detection in the overlying aquifer above the storage reservoir. It helps to quantify the value of individual seismic attributes at multiple monitoring periods based upon their sensitivities.

Kumar, Abhash↗

Computing Nonlinear Power Spectra Across Dynamical Dark Energy Model Space with Neural ODEs

I show how to compute the nonlinear power spectrum across the entire $w(z)$ dynamical dark energy model space. Using synthetic ΛCDM data, I train a neural ordinary differential equation (ODE) to infer the evolution of the nonlinear matter power spectrum as a function of the background expansion and mean matter density across ∼9 Gyr of cosmic evolution. After training, the model generalises to any dynamical dark energy model parameterised by $w(z)$. With little optimisation, the neural ODE is accurate to within 4% up to $k = 5\, h\, {\mathrm Mpc}^{−1}$. Unlike simulation rescaling methods, neural ODEs naturally extend to summary statistics beyond the power spectrum that are sensitive to the growth history.

cosmology↗

A Physics-Based Digital Twin for Wave Elevation and Seabed Moment Estimation of Offshore Monopiles: Preprint

In this work, we present a proof of concept of a physics-based digital twin for a monopile structure (with overhead inertia) subjected to wave loading. The digital twin is formulated using reduced-order models derived from first principles and combined with a Kalman filter for state estimation. The proposed framework estimates the monopile top motion, the wave elevation, and the section forces and moments along the pile using primarily acceleration measurements at the monopile top. Key innovations include the use of a hydrodynamic shape function to represent distributed wave loading in a compact and computationally efficient manner, and the introduction of a shaping filter to augment the state-space with wave kinematics. Synthetic measurement data are generated using OpenFAST and used as a reference to assess the performance of the digital twin. Results demonstrate that the wave elevation can be accurately reconstructed without direct sea-state measurements as long as the wave regime is inertia-dominated. Under the ideal tested conditions, the total hydrodynamic force and sea-bed bending moment are estimated with relative errors on the order of 1% and correlation coefficients exceeding 96%. Future work will evaluate the estimator's performance under operational uncertainties and more complex loading conditions.

17 WIND ENERGY↗

Preliminary Results on Bayesian Inverse UQ for OECD/NEA WPNCS Subgroup 14 Benchmark Exercise for Error Recovery and Experimental Coverage

The Organization for Economic Cooperation and Development (OECD) Working Party on Nucelar Criticality Safety (WPNCS) has proposed a benchmark exercise representative of neutronic behavior in criticality experiments. Here, the goal is to develop confidence in data assimilation techniques used to adjust nuclear data. Participants are given synthetic experimental models with associated measured data and asked to estimate the model parameters given the model and measurements as well as provide predictions for separate application models. In this work, we performed data assimilation using Bayesian inverse Uncertainty Quantification (UQ) with machine learning surrogate models to produce posterior parameter distributions for the requested parameters and posterior predictive distributions for the requested responses. Several experimental models are shown to insufficiently inform the posterior parameter distributions for the applications involved. However, given sufficient experimental data, posterior parameter estimates yielded reduced uncertainty in the response predictions of interest while covering the experimental data.

Bayesian Inference↗

RTN-125: Photometric Transformation Relations for the LSST Data Preview 2

This technical note provides photometric transformation relations between the NSF-DOE Vera C. Rubin Observatory's Data Preview 2 (DP2) and other photometric systems. These transformations are derived using both synthetic and empirical data and are intended to support calibration and comparison across survey systems. We present both polynomial equations and lookup-table-based methods, depending on the available data and desired accuracy. The transformations are generally valid for stars with typical spectral energy distributions (SEDs), and caution should be used when applying them to objects with strong emission lines or atypical colors.

79 ASTRONOMY AND ASTROPHYSICS↗

Summary of the 5th IAEA technical meeting on fusion data processing, validation and analysis (FDPVA)

The purpose of the 5th International Atomic Energy Agency technical meeting on fusion data processing, validation and analysis (FDPVA) (Ghent University, Ghent, Belgium, 12–15 June 2023) was to provide a platform during which a set of topics relevant to FDPVA were discussed with the view of meeting the needs of next step fusion devices such as ITER. The validation and analysis of experimental data obtained from diagnostics used to characterize fusion plasmas are crucial for a knowledge-based understanding of the physical processes governing the dynamics of these plasmas. This paper presents the recent progress and achievements in the domain of plasma diagnostics data analysis and synthetic diagnostics reported at the meeting, including concept description of new devices; fusion databases; integrated data analysis; inverse problems; uncertainty propagation, verification and validation; probabilistic methods and machine learning. The relevant results underline trends observed in the current major fusion confinement devices.

fusion databases↗

RTN-099: Photometric Transformation Relations for the LSST Data Preview 1

This technical note provides photometric transformation relations between the Vera C. Rubin Observatory's LSSTCam and LSSTComCam systems and other photometric systems. These transformations are derived using both synthetic and empirical data and are intended to support calibration and comparison across survey systems. We present both polynomial equations and lookup-table-based methods, depending on the available data and desired accuracy. The transformations are generally valid for stars with typical spectral energy distributions (SEDs), and caution should be used when applying them to objects with strong emission lines or atypical colors.

79 ASTRONOMY AND ASTROPHYSICS↗

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE↗

Improving Runtime Performance of Tensor Computations using Rust From Python

In this work, we investigate improving the runtime performance of key computational kernels in the Python Tensor Toolbox (pyttb), a package for analyzing tensor data across a wide variety of applications. Recent runtime performance improvements have been demonstrated using Rust, a compiled language, from Python via extension modules leveraging the Python C API—e.g., web applications, data parsing, data validation, etc. Using this same approach, we study the runtime performance of key tensor kernels of increasing complexity, from simple kernels involving sums of products over data accessed through single and nested loops to more advanced tensor multiplication kernels that are key in low-rank tensor decomposition and tensor regression algorithms. In numerical experiments involving synthetically generated tensor data of various sizes and these tensor kernels, we demonstrate consistent improvements in runtime performance when using Rust from Python over 1) using Python alone, 2) using Python and the Numba just-in-time Python compiler (for loop-based kernels), and 3) using the NumPy Python package for scientific computing (for pyttb kernels).

97 MATHEMATICS AND COMPUTING↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗