Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Sample Analysis at Mars Instrument Simulator

The Sample Analysis at Mars Instrument Simulator (SAMSIM) is a numerical model dedicated to plan and validate operations of the Sample Analysis at Mars (SAM) instrument on the surface of Mars. The SAM instrument suite, currently operating on the Mars Science Laboratory (MSL), is an analytical laboratory designed to investigate the chemical and isotopic composition of the atmosphere and volatiles extracted from solid samples. SAMSIM was developed using Matlab and Simulink libraries of MathWorks Inc. to provide MSL mission planners with accurate predictions of the instrument electrical, thermal, mechanical, and fluid responses to scripted commands. This tool is a first example of a multi-purpose, full-scale numerical modeling of a flight instrument with the purpose of supplementing or even eliminating entirely the need for a hardware engineer model during instrument development and operation. SAMSIM simulates the complex interactions that occur between the instrument Command and Data Handling unit (C&DH) and all subsystems during the execution of experiment sequences. A typical SAM experiment takes many hours to complete and involves hundreds of components. During the simulation, the electrical, mechanical, thermal, and gas dynamics states of each hardware component are accurately modeled and propagated within the simulation environment at faster than real time. This allows the simulation, in just a few minutes, of experiment sequences that takes many hours to execute on the real instrument. The SAMSIM model is divided into five distinct but interacting modules: software, mechanical, thermal, gas flow, and electrical modules. The software module simulates the instrument C&DH by executing a customized version of the instrument flight software in a Matlab environment. The inputs and outputs to this synthetic C&DH are mapped to virtual sensors and command lines that mimic in their structure and connectivity the layout of the instrument harnesses. This module executes, and thus validates, complex command scripts prior to their up-linking to the SAM instrument. As an output, this module generates synthetic data and message logs at a rate that is similar to the actual instrument.

Benna, Mehdi↗

Measuring and Modeling Shared Visual Attention

Multi-person teams are sometimes responsible for critical tasks, such as flying an airliner. Here we present a method using gaze tracking data to assess shared visual attention, a term we use to describe the situation where team members are attending to a common set of elements in the environment. Gaze data are quantized with respect to a set of N areas of interest (AOIs); these are then used to construct a time series of N dimensional vectors, with each vector component representing one of the AOIs, all set to 0 except for the component corresponding to the currently fixated AOI, which is set to 1. The resulting sequence of vectors can be averaged in time, with the result that each vector component represents the proportion of time that the corresponding AOI was fixated within the given time interval.We present two methods for comparing sequences of this sort, one based on computing the time varying correlation of the averaged vectors, and another based on a chi-square test testing the hypothesis that the observed gaze proportions are drawn from identical probability distributions.We have evaluated the method using synthetic data sets, in which the behavior was modeled as a series of activities, each of which was modeled as a first-order Markov process. By tabulating distributions for pairs of identical and disparate activities, we are able to perform a receiver operating characteristic (ROC) analysis, allowing us to choose appropriate criteria and estimate error rates. Using these criteria, we have applied the methods to data from airline crews, collected in a high-fidelity flight simulator (Gontar Hoermann, 2014). We conclude by considering the problem of automatic (blind) discovery of activities, using methods developed for text analysis.

Mulligan, Jeffrey B.↗

Measuring and Modeling Shared Visual Attention

Multi-person teams are sometimes responsible for critical tasks, such as flying an airliner. Here we present a method using gaze tracking data to assess shared visual attention, a term we use to describe the situation where team members are attending to a common set of elements in the environment. Gaze data are quantized with respect to a set of N areas of interest (AOIs); these are then used to construct a time series of N dimensional vectors, with each vector component representing one of the AOIs, all set to 0 except for the component corresponding to the currently fixated AOI, which is set to 1. The resulting sequence of vectors can be averaged in time, with the result that each vector component represents the proportion of time that the corresponding AOI was fixated within the given time interval. We present two methods for comparing sequences of this sort, one based on computing the time-varying correlation of the averaged vectors, and another based on a chi-square test testing the hypothesis that the observed gaze proportions are drawn from identical probability distributions.We have evaluated the method using synthetic data sets, in which the behavior was modeled as a series of activities, each of which was modeled as a first-order Markov process. By tabulating distributions for pairs of identical and disparate activities, we are able to perform a receiver operating characteristic (ROC) analysis, allowing us to choose appropriate criteria and estimate error rates.We have applied the methods to data from airline crews, collected in a high-fidelity flight simulator (Haslbeck, Gontar Schubert, 2014). We conclude by considering the problem of automatic (blind) discovery of activities, using methods developed for text analysis.

attention↗

Measuring and Modeling Shared Visual Attention

Multi-person teams are sometimes responsible for critical tasks, such as flying an airliner. Here we present a method using gaze tracking data to assess shared visual attention, a term we use to describe the situation where team members are attending to a common set of elements in the environment. Gaze data are quantized with respect to a set of N areas of interest (AOIs); these are then used to construct a time series of N dimensional vectors, with each vector component representing one of the AOIs, all set to 0 except for the component corresponding to the currently fixated AOI, which is set to 1. The resulting sequence of vectors can be averaged in time, with the result that each vector component represents the proportion of time that the corresponding AOI was fixated within the given time interval. We present two methods for comparing sequences of this sort, one based on computing the time-varying correlation of the averaged vectors, and another based on a chi-square test testing the hypothesis that the observed gaze proportions are drawn from identical probability distributions. We have evaluated the method using synthetic data sets, in which the behavior was modeled as a series of "activities," each of which was modeled as a first-order Markov process. By tabulating distributions for pairs of identical and disparate activities, we are able to perform a receiver operating characteristic (ROC) analysis, allowing us to choose appropriate criteria and estimate error rates. We have applied the methods to data from airline crews, collected in a high-fidelity flight simulator (Haslbeck, Gontar & Schubert, 2014). We conclude by considering the problem of automatic (blind) discovery of activities, using methods developed for text analysis.

attention↗

Surface Reflectance of Mars Observed by CRISM-MRO: 1. Multi-angle Approach for Retrieval of Surface Reflectance from CRISM Observations (mars-reco)

This article addresses the correction for aerosol effects in near-simultaneous multiangle observations acquired by the Compact Reconnaissance Imaging Spectrometer for Mars (CRISM) aboard the Mars Reconnaissance Orbiter. In the targeted mode, CRISM senses the surface of Mars using 11 viewing angles, which allow it to provide unique information on the scattering properties of surface materials. In order to retrieve these data, however, appropriate strategies must be used to compensate the signal sensed by CRISM for aerosol contribution. This correction is particularly challenging as the photometric curve of these suspended particles is often correlated with the also anisotropic photometric curve of materials at the surface. This article puts forward an innovative radiative transfer based method named Multi-angle Approach for Retrieval of Surface Reflectance from CRISM Observations (MARS-ReCO). The proposed method retrieves photometric curves of surface materials in reflectance units after removing aerosol contribution. MARS-ReCO represents a substantial improvement regarding previous techniques as it takes into consideration the anisotropy of the surface, thus providing more realistic surface products. Furthermore, MARS-ReCO is fast and provides error bars on the retrieved surface reflectance. The validity and accuracy of MARS-ReCO is explored in a sensitivity analysis based on realistic synthetic data. According to experiments, MARS-ReCO provides accurate results (up to 10 reflectance error) under favorable acquisition conditions. In the companion article, photometric properties of Martian materials are retrieved using MARS-ReCO and validated using in situ measurements acquired during the Mars Exploration Rovers mission.

CRISM/MRO↗

A comparison of effective field theory models of redshift space galaxy power spectra for DESI 2024 and future surveys

In preparation for the next generation of galaxy redshift surveys, and in particular the year-one data release from the Dark Energy Spectroscopic Instrument (DESI), we investigate the consistency of a variety of effective field theory models that describe the galaxy-galaxy power spectra in redshift space into the quasi-linear regime using 1-loop perturbation theory. These models are employed in the pipelines velocileptors, PyBird, and Folpsν. While these models have been validated independently, a detailed comparison with consistent choices has not been attempted. After briefly discussing the theoretical differences between the models we describe how to provide a more apples-to-apples comparison between them. We present the results of fitting mock spectra from the AbacusSummit suite of N-body simulations provided in three redshift bins to mimic the types of dark time tracers targeted by the DESI survey. We show that the theories behave similarly and give consistent constraints in both the forward-modeling and ShapeFit compressed fitting approaches. We additionally generate (noiseless) synthetic data from each pipeline to be fit by the others, varying the scale cuts in order to show that the models agree within the range of scales for which we expect 1-loop perturbation theory to be applicable. Finally, this work lays the foundation of Full-Shape analysis with DESI Y1 galaxy samples where in the tests we performed, we found no systematic error associated with the modeling of the galaxy redshift space power spectrum for this volume.

79 ASTRONOMY AND ASTROPHYSICS↗

Adaptive Data Screening for Multi-Angle Polarimetric Aerosol and Ocean Color Remote Sensing Accelerated by Automatic Differentiation

Remote sensing measurements from multi-angle polarimeters (MAPs) contain rich aerosol microphysical property information, and these sensors have been used to perform retrievals in optically complex atmosphere and ocean systems. Previous studies have concluded that, generally, five moderately separated viewing angles in each spectral band provide sufficient accuracy for aerosol property retrievals, with performance gradually saturating as angles are added above that threshold. The Hyper-Angular Rainbow Polarimeter (HARP) instruments provide high angular sampling with a total of 90-120 unique angles across four bands, a capability developed mainly for liquid cloud retrievals. In practice, not all view angles are optimal for aerosol retrievals due to impacts of clouds, sun glint, and other impediments. The many viewing angles of HARP can provide resilience to these effects, if the impacted views are screened from the dataset, as the remaining views may be sufficient for successful analysis. In this study, we discuss how the number of available viewing angles impacts aerosol and ocean color retrieval uncertainties, as applied to two versions of the HARP instrument. AirHARP is an airborne prototype that was deployed in the ACEPOL field campaign, while HARP2 is an instrument in development for the upcoming NASA Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission. Based on synthetic data, we find that a total of 20-30 angles across all bands (i.e. five to eight viewing angles per band) are sufficient to achieve good retrieval performance. Following from this result, we develop an adaptive multi-angle polarimetric data screening (MAPDS) approach to evaluate data quality by comparing measurements with their best-fitted forward model. The FastMAPOL retrieval algorithm is used to retrieve scene geophysical values, by matching an efficient, deep learning-based, radiative transfer emulator to observations. The data screening method effectively identifies and removes viewing angles affected by thin cirrus clouds and other anomalies, improving retrieval performance. This was tested with AirHARP data, and we found agreement with the High Spectral Resolution Lidar-2 (HSRL-2) aerosol data. The data screening approach can be applied to modern satellite remote sensing missions, such as PACE, where a large amount of multi-angle, hyperspectral, polarimetric measurements will be collected.

multi-angle polarimeter↗

Neural network-based classification and regression of magnetohydrodynamic modes in tokamaks

We present a machine learning-based magnetohydrodynamic (MHD) classifier and regressor that utilizes real or complex-valued 3D magnetic sensor array data to determine neoclassical tearing mode (NTM) onset times in tokamaks with millisecond accuracy. The input dataset consists of poloidal profiles of complex Fourier amplitudes with an n = 1 toroidal mode number from 144 human-labeled ITER Baseline Scenario discharges in the DIII-D tokamak, spanning both tearing-dominated and sawtooth-dominated regimes. Since m, n = 2,1 NTMs frequently emerge alongside sawteeth at the same frequency in this scenario, the focus is on isolating the m = 1 and m = 2 components of the n = 1 MHD mode near the tearing onset. To improve model regularization and prediction stability, singular value decomposition was applied to balance the sawtooth and tearing datasets. The enriched datasets facilitated training neural networks that learn the key distinguishing features of sawtooth and tearing modes in the poloidal profiles of their magnetic amplitude and phase. When the modes occur independently, the networks achieve perfect classification due to the modes’ distinct characteristics and low measurement noise. In the more experimentally relevant case where both modes coexist, the networks maintain exceptional performance across key metrics. Tests on synthetic data with known ground truth demonstrate the superior accuracy of the neural network trained on complex-valued input compared to models using real amplitude, phase, or pseudo-complex data, achieving both a mean time delay and standard deviation below 1 ms. Notably, standard linear regression methods fitting the dominant singular modes to the data closely match the neural network’s performance. Applying these methods across a broad range of H-mode scenarios will enable future studies to systematically identify dominant NTM triggers as scenario-specific variables, paving the way for more effective tearing mode avoidance strategies in future fusion reactor designs.

machine learning↗

Denoising diffusion probabilistic models for generative alloy design

Inverse material design is an extremely challenging optimization task made difficult by, in part, the highly nonlinear relationship linking performance with composition. Quantitative approaches have improved significantly owing to advances in high throughput experimentation and computational thermodynamics. However, existing physics-based tools are mostly forward models; input a chemistry and obtain a prediction. More recently the materials community has leveraged advances in the machine learning community to establish novel inverse design frameworks. Very recently denoising diffusion probabilistic models have been shown to be extremely powerful generators producing synthetic data of various modalities e.g. images, text, audio, tables, etc.. In this work a novel framework for alloy design and optimization is proposed leveraging these class of models. Five key generative tasks are demonstrated (1) unconditional generation (2) composition conditioned generation (3) property conditioned generation (4) multi-feedstock conditioned generation and (5) generative optimization. These methods were tested on three case studies: high entropy alloy design, superalloy binder jet additive manufacturing, and in-situ dual-feedstock wire-arc additive manufacturing. Results indicate that the established models are extremely flexible, expressive, and robust. The architecture’s flexibility and training procedure empower the model to learn complex intra-compositional and composition-property relationships. Furthermore, the probabilistic nature of these models makes them well suited for addressing solution non-uniqueness and tackling uncertainty quantification tasks. While the fidelity and quantity of the underlying training data is paramount, we envision that future alloy design frameworks will make extensive use of these kinds of machine learning models as “search” tools bolstering the utility of experimental and computational approaches.

36 MATERIALS SCIENCE↗

Predicting critical heat flux with uncertainty quantification and domain generalization using conditional variational autoencoders and deep neural networks

Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evolution and Degradation Patterns of Electrochemical Cells Based on the Analysis of Interfacial Phenomena at Li Metal Anode/Electrolyte Interfaces

In this work, we report the results of a theoretical–computational analysis of the solid electrolyte interphase (SEI) growth and degradation dynamics occurring in lithium metal batteries during cycling. We use ab initio-kinetic Monte Carlo simulations to generate a synthetic data set, which is analyzed by machine learning methods. We aim to determine: (i) how modifications in interfacial interaction energies between solid electrolyte interphase (SEI) blocks and between Li ions and SEI facets impact the Coulombic efficiency (CE) of the battery and (ii) what factors, including reactions, microscopic transport, and other interfacial events, may lead to cell performance “failure” during prolonged charge and discharge cycles, signaled as a sharp decay in the CE over cycling. The demonstration of our approach is done on a cell including a Li metal surface interfacing with a previously introduced state-of-the-art electrolyte, and the idea can be applied to any electrochemical system. Outcomes include the identification of the leading chemical, physical, and structural variables causing cell failure and relating them to the electrolyte formulation, thus paving the way to future more refined analysis and electrolyte design.

batteries↗

A Framework for Parametric and Predictive Uncertainty Quantification in the E3SM Land Model: Assessing Site and Observable Generalizability

Quantifying parametric uncertainty using observations from individual sites provides a critical foundation for Earth system modeling, serving as a necessary first step before scaling up to regional or global applications. This study introduces a novel computational framework designed to enhance model predictability by reducing parametric uncertainty and assessing site and observable generalizability using various observational constraints. The framework integrates five components: Model Simulation, Statistical Emulation, Global Sensitivity Analysis (GSA), Model Calibration, and Model Prediction. Using the E3SM land model, we simulated site-level land-atmosphere carbon and energy fluxes from 2003 to 2007 across five evergreen needleleaf FLUXNET sites, perturbing 26 vegetation-related model parameters. Gaussian process emulators were employed to expedite GSA and model calibration. Four critical parameters that strongly influence selected land-atmosphere fluxes were identified by GSA. Bayesian approaches were used to infer parameter probability distributions leveraging synthetic data and FLUXNET observations. The results reveal that posterior parameter distributions vary significantly across different sites and observables within the same plant functional type. Probabilistic predictions indicate that parameters calibrated at one site can enhance predictive accuracy at other sites, although site heterogeneity may sometimes outweigh parametric uncertainty. Additionally, the probabilistic predictions demonstrate that calibration for one variable can also improve predictability for other variables, thereby maximizing predictive capabilities with limited observations. This framework provides a powerful approach for reducing parametric uncertainty in Earth system models and deepening our understanding of carbon dynamics and energy cycles. Its adaptability makes it a valuable tool for broader applications in Earth system modeling.

54 ENVIRONMENTAL SCIENCES↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

Source localization for neutron imaging systems using convolutional neural networks

The nuclear imaging system at the National Ignition Facility (NIF) is a crucial diagnostic for determining the geometry of inertial confinement fusion implosions. The geometry is reconstructed from a neutron aperture image via a set of reconstruction algorithms using an iterative Bayesian inference approach. An important step in these reconstruction algorithms is finding the fusion source location within the camera field-of-view. Currently, source localization is achieved via an iterative optimization algorithm. In this paper, we introduce a machine learning approach for source localization. Specifically, we train a convolutional neural network to predict source locations given a neutron aperture image. We show that this approach decreases computation time by several orders of magnitude compared to the current optimization-based source localization while achieving similar accuracy on both synthetic data and a collection of recent NIF deuterium–tritium shots.

47 OTHER INSTRUMENTATION↗

Inferring performance metrics for laser direct drive experiments on OMEGA

Quantifying performance improvements on the OMEGA laser facility requires robust inference of established no-alpha performance metrics, which requires, at minimum, a model to infer the shocked fuel mass and pressure of the confined fusion plasma. In this work, we describe the methodology used to infer performance metrics on OMEGA and present the current state-of-the art model used to infer these metrics from OMEGA experiments. In particular, since neutron images of cryogenic implosions are not available on OMEGA at present, we present how x-ray sizes are determined on OMEGA using a Gaussian Process regression model and how the neutron production region's size is inferred from them. As a result, we end by benchmarking the model using synthetic data and 1-D LILAC simulations and test its experimental self-consistency across available x-ray diagnostic channels.

Gopalaswamy, V. [Laboratory for Laser Energetics, ↗

From minimum-viable-products to full models: a step-wise development of diagnostic forward models in support of design, analysis and modelling on the ST40 tokamak

Like most magnetic confined fusion experiments, the ST40 tokamak started off with a small subset of diagnostics and gradually increased the diagnostic set to include more complex and comprehensive systems. To make the most of each operational phase, forward models of various diagnostics are used and developed to aid design, provide consistency-checks during commissioning, test analysis methods, and build workflows to constrain high-level parameters to inform interpretation, theory and modelling. For new models and new analysis workflows, minimum-viable-products are released early, and their complexity is increased in a step-wise manner, facilitating the support of all programme phases on multiple parallel applications, while enabling learning opportunities and feedback loops. In this contribution we review the philosophy, scope and architecture of the framework under development. We discuss the details of some forward models, with examples on how they are used to aid diagnostic design, to investigate analysis methodologies through synthetic data, and how they are embedded in experimental analysis workflows. We compare previously published experimental results with new, more advanced analysis workflows employing more recent, detailed models and new diagnostic data, providing confirmation of the published material from the 2021–22 experimental campaign.

integrated data analysis↗

Statistical inference of anomalous thermal transport with uncertainty quantification for interpretive 2D SOL models

The critical task of inferring anomalous cross-field transport coefficients is addressed in simulations of boundary plasmas with fluid models. A workflow for parameter inference in the UEDGE fluid code is developed using Bayesian optimization with parallelized sampling and integrated uncertainty quantification. In this workflow, transport coefficients are inferred by maximizing their posterior probability distribution, which is generally multidimensional and non-Gaussian. Uncertainty quantification is integrated throughout the optimization within the Bayesian framework that combines diagnostic uncertainties and model limitations. As a concrete example, we infer the anomalous electron thermal diffusivity $\chi_\perp$ from an interpretive 2D model describing electron heat transport in the conduction-limited region with radiative power loss. The workflow is first benchmarked against synthetic data and then tested on H-, L-, and I-mode discharges to match their midplane temperature and divertor heat flux profiles. We demonstrate that the workflow efficiently infers diffusivity and its associated uncertainty, generating 2D profiles that match 1D measurements. Future efforts will focus on incorporating more complicated fluid models and analyzing transport coefficients inferred from a large database of experimental results.

Bayesian optimization↗