Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Machine learning based unfolding of x-ray spectra from filter stack spectrometer data

We demonstrate the application of neural networks to perform x-ray spectra unfolding from data collected by filter stack spectrometers. A filter stack spectrometer consists of a series of filter-detector pairs, where the detectors behind each filter measure the energy deposition through each layer as photo-stimulated luminescence (PSL). The network is trained on synthetic data, assuming x-rays of energies < 1 MeV and of two different distribution functions (Maxwellian and Gaussian) and the corresponding measured PSL values obtained from five different filter stack spectrometer designs. Predicted unfolds of single distributions are near identical reproductions of the ground truth spectra, with differences in the values lower than 20% at the higher energy end in some cases. The neural network has also demonstrated robustness to experimental measurement errors of < 5% and some capability of performing unfolds for linear combinations of the two distributions without previous training. The network can perform unfolds at rates > 1 Hz, ideal for application to some high-repetition-rate systems.

47 OTHER INSTRUMENTATION↗

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING↗

Characterization of contaminants in the Lyman-alpha forest auto-correlation with DESI

Baryon Acoustic Oscillations can be measured with sub-percent precision above redshift two with the Lyman-α (Lyα) forest auto-correlation and its cross-correlation with quasar positions. This is one of the key goals of the Dark Energy Spectroscopic Instrument (DESI) which started its main survey in May 2021. We present in this paper a study of the contaminants to the Lyα forest which are mainly caused by correlated signals introduced by the spectroscopic data processing pipeline as well as astrophysical contaminants due to foreground absorption in the intergalactic medium. Notably, an excess signal caused by the sky background subtraction noise is present in the Lyα auto-correlation in the first line-of-sight separation bin. We use synthetic data to isolate this contribution, we also characterize the effect of spectro-photometric calibration noise, and propose a simple model to account for both effects in the analysis of the Lyα forest. We then measure the auto-correlation of the quasar flux transmission fraction of low redshift quasars, where there is no Lyα forest absorption but only its contaminants. We demonstrate that we can interpret the data with a two-component model: data processing noise and triply ionized Silicon and Carbon auto-correlations. This result can be used to improve the modeling of the Lyα auto-correlation function measured with DESI.

79 ASTRONOMY AND ASTROPHYSICS↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Testing Fractal Methods on Observed and Simulated Solar Magnetograms

The term "magnetic complexity" has not been sufficiently quantified. To accomplish this, we must understand the relationship between the observed magnetic field of solar active regions and fractal dimension measurements. Using data from the Marshall Space Flight Center's vector magnetograph ranging from December 1991 to July 2001, we compare the results of several methods of calculating a fractal dimension, e.g., Hurst coefficient, the Higuchi method, power spectrum, and 2-D Wavelet Packet Analysis. In addition, we apply these methods to synthetic data, beginning with representations of very simple dipole regions, ending with regions that are magnetically complex.

Adams, M.↗

Limited Angle Reconstruction Method for Reconstructing Terrestrial Plasmaspheric Densities from EUV Images

A new method for reconstructing the global 3D distribution of plasma densities in the plasmasphere from a limited number of 2D views is presented. The method is aimed at using data from the Extreme Ultra Violet (EUV) sensor on NASA s Imager for Magnetopause-to-Aurora Global Exploration (IMAGE) satellite. Physical properties of the plasmasphere are exploited by the method to reduce the level of inaccuracy imposed by the limited number of views. The utility of the method is demonstrated on synthetic data.

Newman, Timothy↗

Chemical Modeling of the Reactivity of Short-Lived Greenhouse Gases: A Model Inter-Comparison Prescribing a Well-Measured, Remote Troposphere

We develop a new protocol for merging in situ measurements with 3-D model simulations of atmospheric chemistry with the goal of integrating over the data to identify the most reactive air parcels in terms of tropospheric production and loss of the greenhouse gases ozone and methane. Presupposing that we can accurately measure atmospheric composition, we examine whether models constrained by such measurements agree on the chemical budgets for ozone and methane. In applying our technique to a synthetic data stream of 14,880 parcels along 180W, we are able to isolate the performance of the photochemical modules operating within their global chemistry-climate and chemistry-transport models, removing the effects of modules controlling tracer transport, emissions, and scavenging. Differences in reactivity across models are driven only by the chemical mechanism and the diurnal cycle of photolysis rates, which are driven in turn by temperature, water vapor, solar zenith angle, clouds, and possibly aerosols and overhead ozone, which are calculated in each model. We evaluate six global models and identify their differences and similarities in simulating the chemistry through a range of innovative diagnostics. All models agree that the more highly reactive parcels dominate the chemistry (e.g., the hottest 10% of parcels control 25-30% of the total reactivities), but do not fully agree on which parcels comprise the top 10%. Distinct differences in specific features occur, including the regions of maximum ozone production and methane loss, as well as in the relationship between photolysis and these reactivities. Unique, possibly aberrant, features are identified for each model, providing a benchmark for photochemical module development. Among the 6 models tested here, 3 are almost indistinguishable based on the inherent variability caused by clouds, and thus we identify 4, effectively distinct, chemical models. Based on this work, we suggest that water vapor differences in model simulations of past and future atmospheres may be a cause of the different evolution of tropospheric O3 and CH4, and lead to different chemistry-climate feedbacks across the models.

greenhouse gases↗

Magsat science investigations

Existing software is being modified to take any combination of component or scalar data in profile form and invert it to a discrete-source magnetization distribution for sources having arbitrary equal-area spacing. The option of constraining both source and directions and magnitude is included. Software for spectral depth-to-magnetic bottom estimates is under development. The software is to be thoroughly listed on synthetic data and applied to the NOO survey data and to NURE data for the southern Rio Grande Rift. Swanberg's silica geotemperature data for the U.S. was digitized for heat flow studies.

Source record↗

Optimized Profile Retrievals of Aerosol Microphysical Properties from Simulated Spaceborne Multiwavelength Lidar

This work is an expanded study of one previously published onretrievals of aerosol microphysical properties from space-borne multiwavelengthlidarmeasurements. The earlier studiesand this one weredone in the framework of the NASA Aerosol-Clouds-Ecosystems (now the Aerosol Clouds Convection and Precipitation) NASA mission. The focus here is on the capabilities of a simulated spacebornemultiwavelengthlidar system for retrieving aerosol complex refractive index (m = mr+ imi) and spectral single scattering albedo (SSA(λ)), although other bulk parameters such as effective (reff) radius and particle volume (V) and surface (S) concentrations are also studied. The novelty presented here is the use of recently published, case dependent optimized-constraints on the microphysical retrievals using three backscattering coefficients (β) at 355, 532 and 1064 nm and two extinction coefficients (α) at 355 and 532 nm, typically known as the stand-alone 3β+2α lidar inversion. Case-dependent optimized-constraints (CDOC) limit the ranges of refractive index, both real (mr) and imaginary (mi) parts, and of radii that are permitted in the retrievals. Such constraints are selected directly from the 3β+2α41measurements through an analysis of the relationship between spectral dependence of aerosol extinction-to-backscatter ratios (LR) and the Ångström exponent of extinction. The analyses presented here for different sets ofsize distributions and refractive indices reveal that the direct determination of CDOCareonly feasible for cases where the uncertaintiesin the input optical data areless than 15 %.Forthe same simulated spacebornesystem and yield than in Whiteman et al., (2018), we demonstrated that the use of CDOC as essential for the retrievals of refractive index and also largely improved retrieval of bulk parameters. A discussion of the global representativeness of CDOC is presented using simulated lidar data from a 24-hour satellite track using GEOS model output to initialize the lidar simulator.We found that CDOCare representative of many aerosol mixtures in spite of some outliers (e.g. highly hydrated particles) associatedwith the assumptions of bimodal size distributions and of the same refractive index for fine and coarse modes. Moreover, sensitivity tests performed using synthetic data reveal that retrievals of imaginary refractive index (mi) and SSA are extremely sensitive to β(355).

NASA Aerosol-Clouds-Ecosystems↗

Dispersive and nondispersive 𝐾-matrix formalisms

The modeling of coupled-channel effects has become increasingly important due to the availability of highly precise data for a large variety of hadronic (re)scattering processes. The 𝐾-matrix is a powerful, yet comparatively simple, method to describe scattering amplitudes, including coupled-channel effects, with the aim of interpreting experimental data. Throughout the literature, a range of dispersive and nondispersive 𝐾-matrix methods are employed. Here, we compare the dispersive and nondispersive formulations in the context of the N/D method. It is shown that the methods are equivalent in the physical region under 𝐾-matrix reparametrization. Differences away from the physical region are examined. Applications to synthetic data are used to illustrate the effects of model choices concerning form factors and the application of dispersion relations, with the goal of clarifying best practices. We find no clear preference with regard to dispersive modeling. In contrast, we find that interpretational ambiguity of the bare model parameters—and even of the form of the bare model—is endemic, and recommend a thorough sampling of data and model spaces to assess conclusion robustness.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Metric to Quantify Shared Visual Attention in Two-Person Teams

Introduction: Critical tasks in high-risk environments are often performed by teams, the members of which must work together efficiently. In some situations, the team members may have to work together to solve a particular problem, while in others it may be better for them to divide the work into separate tasks that can be completed in parallel. We hypothesize that these two team strategies can be differentiated on the basis of shared visual attention, measured by gaze tracking. 2) Methods: Gaze recordings were obtained for two-person flight crews flying a high-fidelity simulator (Gontar, Hoermann, 2014). Gaze was categorized with respect to 12 areas of interest (AOIs). We used these data to construct time series of 12 dimensional vectors, with each vector component representing one of the AOIs. At each time step, each vector component was set to 0, except for the one corresponding to the currently fixated AOI, which was set to 1. This time series could then be averaged in time, with the averaging window time (t) as a variable parameter. For example, when we average with a t of one minute, each vector component represents the proportion of time that the corresponding AOI was fixated within the corresponding one minute interval. We then computed the Pearson product-moment correlation coefficient between the gaze proportion vectors for each of the two crew members, at each point in time, resulting in a signal representing the time-varying correlation between gaze behaviors. We determined criteria for concluding correlated gaze behavior using two methods: first, a permutation test was applied to the subjects' data. When one crew member's gaze proportion vector is correlated with a random time sample from the other crewmember's data, a distribution of correlation values is obtained that differs markedly from the distribution obtained from temporally aligned samples. In addition to validating that the gaze tracker was functioning reasonably well, this also allows us to compute probabilities of coordinated behavior for each value of the correlation. As an alternative, we also tabulated distributions of correlation coefficients for synthetic data sets, in which the behavior was modeled as a first-order Markov process, and compared correlation distributions for identical processes with those for disparate processes, allowing us to choose criteria and estimate error rates. 3) Discussion: Our method of gaze correlation is able to measure shared visual attention, and can distinguish between activities involving different instruments. We plan to analyze whether pilots strategies of sharing visual attention can predict performance. Possible measurements of performance include expert ratings from instructors, fuel consumption, total task time, and failure rate. While developed for two-person crews, our approach can be applied to larger groups, using intra-class correlation coefficients instead of the Pearson product-moment correlation.

Gontar, Patrick↗

Analysis of data acquired by synthetic aperture radar and LANDSAT Multispectral Scanner over Kershaw County, South Carolina, during the summer season

Data acquired by synthetic aperture radar (SAR) and LANDSAT multispectral scanner (MSS) were processed and analyzed to derive forest-related resources inventory information. The SAR data were acquired by using the NASA aircraft X-band SAR with linear (HH, VV) and cross (HV, VH) polarizations and the SEASAT L-band SAR. After data processing and data quality examination, the three polarization (HH, HV, and VV) data from the aircraft X-band SAR were used in conjunction with LANDSAT MSS for multisensor data classification. The results of accuracy evaluation for the SAR, MSS and SAR/MSS data using supervised classification show that the SAR-only data set contains low classification accuracy for several land cover classes. However, the SAR/MSS data show that significant improvement in classification accuracy is obtained for all eight land cover classes. These results suggest the usefulness of using combined SAR/MSS data for forest-related cover mapping. The SAR data also detect several small special surface features that are not detectable by MSS data.

Wu, S. T.↗

A physical model for predicting bidirectional reflectances over bare soil

While most previous attempts to retrieve soil surface optical property characteristics have proceeded through a fitting of empirical functions to the data, an optimization technique is presently applied to a physically-based surface reflectance model developed for the study of planetary surfaces. This inversion procedure is shown to allow the direct estimation of the single-scattering coefficient, two parameters describing the 'hot spot' phenomenon, and two parameters describing the scattering phase function. A comparison of inversion technique results with both synthetic data and actual observations shows the model to be capable of predicting the observed bidirectional reflectances as well as directional-hemispherical reflectances; it can also build the complete radiance field over the upward hemisphere.

Pinty, Bernard↗

Bond Line Thickness Estimation in Composite Structures Using Multiple Inspection Techniques

Imaging and other nondestructive evaluation techniques are commonly used for material characterization and defect recognition in safety critical aerospace applications, with data fusion providing the framework for uncertainty quantification in these contexts. Most commonly, forward physics-based modeling predicts the response conditioned on material properties and defect assumptions, and probabilistic methods are used to infer the hidden state of the subject of the inspection from a combination of prior information, likelihoods, and inspection data. In this paper Bayesian methods are used to estimate bond thickness in lap joints comprised of aluminum adherends using a combination of infrared thermography and ultrasound. The concept of the conflation of probability distributions is applied to combine the posterior distributions derived from thermography and ultrasound and the quality of the fused estimates are compared against the individual estimates against synthetic data that was created to mimic the inspection of a lap joint comprised of aluminum adherends.

thermal nondestructive evaluation↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a pilot flying an airplane; in this case, activities with distinct visual signatures might be things like communicating, navigating, monitoring, etc. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. We compare, not individual fixations, but groups of fixations aggregated over a fixed time interval (Tau). We assume that the environment has been divided into a finite set of discrete areas-of-interest (AOIs). For a given time interval, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector, where N is the number of AOIs. These proportions can be converted to integer counts by multiplying by Tau divided by the average fixation duration, a parameter that we fix at 283 milliseconds. We compare different intervals by computing the chi-squared statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We cluster the intervals, first by merging adjacent intervals that are sufficiently similar, optionally shifting the boundary between non-merged intervals to maximize the difference. Then we compare and cluster non-adjacent intervals. The method is evaluated using synthetic data generated by a hand-crafted set of activities. While the method generally finds more activities than put into the simulation, we have obtained agreement as high as 80 percent between the inferred activity labels and ground truth.

Eye Movements↗

Regularizing INR with Diffusion Prior for Self-Supervised 3D Reconstruction OF Neutron Computed Tomography Data

Recently, generative diffusion priors have made huge strides as inverse problem solvers, including the ability to be adapted for inference on out-of-distribution data. Concurrently, implicit neural representations (INRs) have emerged as fast and lightweight inverse imaging solvers that are amenable to hybrid approaches that combine learned priors with traditional inverse problem formulations. In this paper, we present a diffusive computed tomography (CT) inversion framework for regularizing INRs called Diffusive INR (DINR), designed to enable high-quality reconstruction from sparse-view neutron CT. Pretrained purely on synthetic data, DINR is evaluated on simulated and experimentally obtained observations of concrete microstructures, where traditional reconstruction methods suffer substantial degradation when the number of views is reduced. Our approach delivers superior performance, reduces reconstruction artifacts, and achieves gains in PSNR and SSIM, enabling accurate micro-structural characterization even under extreme data limitations compared to state-of-the-art sparse-view reconstruction techniques.

Hossain, Maliha [ORNL]↗