Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

Spectral Information from Gapped Data. a Comparison of Technique

The fast Fourier transformations (FFT) is used to estimate power spectra of continuous signals evenly sampled on discrete domains. The problem of finding power spectra on unevenly sampled domains, in particular a regularly spaced domain with gaps is discussed. The analysis of the ACRIM solar bolometric intensity data, obtained with a 3/5 on and 2/5 off duty cycle of approximately 100 minutes, would benefit from the techniques. The comparative effectiveness of three different analysis techniques applied to synthetic data generated on gapped domain is reported.

Kuhn, J. R.↗

Software for Simulating Remote Sensing Systems

The Application Research Toolbox (ART) is a collection of computer programs that implement algorithms and mathematical models for simulating remote sensing systems. The ART is intended to be especially useful for performing design-tradeoff studies and statistical analyses to support the rational development of design requirements for multispectral imaging systems. Among other things, the ART affords a capability to synthesize coarser-spatial-resolution image-data sets from finer-spatial-resolution data sets and multispectral-image-data products from hyperspectral-image-data products. The ART also provides for synthesis of image-degradation effects, including point-spread functions, misregistration of spectral images, and noise. The ART can utilize real or synthetic data sets, along with sensor specifications, to create simulated data sets. In one example of a typical application, simulated data pertaining to an existing multispectral sensor system are used to verify the data collected by the system in operation. In the case of a proposed sensor system, the simulated data can be used to conduct trade studies and statistical analyses to ensure that the sensor system will satisfy the requirements of potential scientific, academic, and commercial user communities.

Zanoni, Vicki↗

Software for Simulating Remote Sensing Systems

The Application Research Toolbox (ART) is a collection of computer programs that implement algorithms and mathematical models for simulating remote sensing systems. The ART is intended to be especially useful for performing design-tradeoff studies and statistical analyses to support the rational development of design requirements for multispectral imaging systems. Among other things, the ART affords a capability to synthesize coarser-spatial-resolution image-data sets from finer-spatial-resolution data sets and multispectral-image-data products from hyperspectral-image-data products. The ART also provides for synthesis of image-degradation effects, including point-spread functions, misregistration of spectral images, and noise. The ART can utilize real or synthetic data sets, along with sensor specifications, to create simulated data sets. In one example of a typical application, simulated data pertaining to an existing multispectral sensor system are used to verify the data collected by the system in operation. In the case of a proposed sensor system, the simulated data can be used to conduct trade studies and statistical analyses to ensure that the sensor system will satisfy the requirements of potential scientific, academic, and commercial user communities.

Vicki Zanoni↗

Simulating Remote Sensing Systems

The Application Research Toolbox (ART) is a collection of computer programs that implement algorithms and mathematical models for simulating remote sensing systems. The ART is intended to be especially useful for performing design-tradeoff studies and statistical analyses to support the rational development of design requirements for multispectral imaging systems. Among other things, the ART affords a capability to synthesize coarser-spatial-resolution image-data products. The ART also provides for simulations of image-degradation effects, including point-spread functions, misregistration of spectral images, and noise. The ART can utilize real or synthetic data sets, along with sensor specifications, to create simulated data sets. In one example of a particular application, simulated imagery of a coarse resolution system was created using high-resolution imagery from another system in order to perform a radiometric cross-comparison. In the case of a proposed sensor system, the simulated data can be used to conduct trade studies and statistical analyses to ensure that the sensor system will satisfy the requirements of potential scientific, academic, and commercial user communities.

Zanoni, Vicki↗

HumoNet: A Framework for Realistic Modeling and Simulation of Human Mobility Network

Understanding, analyzing, and predicting human mobility and dynamics are valuable to solving pressing problems, developing effective plans, and prescribing timely remedies. As a computational approach, realistic human mobility simulations allow us to understand, analyze, and predict complex systems, including human societies. Accurate simulations rely on (1) the model that captures interactions and behaviors of myriad entities in our society and (2) the mapping of model instances to real-world entities. Taking this into account, this paper introduces the Human Mobility Network simulation framework (HumoNet), an integrated patterns of life (POL) simulation framework that leverages real-world data layers including transportation networks, points of interest, populations, popularity, and human trajectories. HumoNet is a data informed model in which agents are equipped with activities, locomotion, and planning capabilities. To simulate realistic kinematic maneuvers of individuals in transportation networks, HumoNet harnesses a microscopic traffic simulator that provides interaction among vehicles and traffic objects. In this paper, we describe the framework, outline our methodologies, and discuss the data processing and challenges of each data layer. Through experiments, we demonstrate that our simulations capture key features of human mobility by comparing them to the literature and real data using standard measures of human mobility (i.e., the radius of gyration, number of locations visited, level of exploration) and metrics scoring (i.e., Jensen-Shannon divergence). We envision that the synthetic data produced by HumoNet will serve as a benchmark for analyzing epidemics, deploying EV charging networks, and validating AI/ML tasks such as location prediction.

Kim, Joon-Seok↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Efficient bulk-loading of gridfiles

This paper considers the problem of bulk-loading large data sets for the gridfile multiattribute indexing technique. We propose a rectilinear partitioning algorithm that heuristically seeks to minimize the size of the gridfile needed to ensure no bucket overflows. Empirical studies on both synthetic data sets and on data sets drawn from computational fluid dynamics applications demonstrate that our algorithm is very efficient, and is able to handle large data sets. In addition, we present an algorithm for bulk-loading data sets too large to fit in main memory. Utilizing a sort of the entire data set it creates a gridfile without incurring any overflows.

Leutenegger, Scott T.↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

Online anomaly detection↗

Distributed Target Tracking With Optimal Data Migration

The paper presents an Extended Kalman Filter based framework for airborne target tracking using dynamic information fusion from multi-modal sensors with geodiversity. First, the algorithm execution location is determined using an optimal data migration strategy, next the sensors information is dynamically fused at each estimation instance using validity flag for each sensor reading, finally the target estimation is updated based on the fused innovation vector. The approach is applied to synthetic data generated from the radar and camera models located on the ground for the simulated target flight in Reflection simulation environment.

Distributed sensing↗

Forest biomass, canopy structure, and species composition relationships with multipolarization L-band synthetic aperture radar data

The effect of forest biomass, canopy structure, and species composition on L-band synthetic aperature radar data at 44 southern Mississippi bottomland hardwood and pine-hardwood forest sites was investigated. Cross-polarization mean digital values for pine forests were significantly correlated with green weight biomass and stand structure. Multiple linear regression with five forest structure variables provided a better integrated measure of canopy roughness and produced highly significant correlation coefficients for hardwood forests using HV/VV ratio only. Differences in biomass levels and canopy structure, including branching patterns and vertical canopy stratification, were important sources of volume scatter affecting multipolarization radar data. Standardized correction techniques and calibration of aircraft data, in addition to development of canopy models, are recommended for future investigations of forest biomass and structure using synthetic aperture radar.

Sader, Steven A.↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Basic Diagnosis and Prediction of Persistent Contrail Occurrence using High-resolution Numerical Weather Analyses/Forecasts and Logistic Regression. Part I: Effects of Random Error

Straightforward application of the Schmidt-Appleman contrail formation criteria to diagnose persistent contrail occurrence from numerical weather prediction data is hindered by significant bias errors in the upper tropospheric humidity. Logistic models of contrail occurrence have been proposed to overcome this problem, but basic questions remain about how random measurement error may affect their accuracy. A set of 5000 synthetic contrail observations is created to study the effects of random error in these probabilistic models. The simulated observations are based on distributions of temperature, humidity, and vertical velocity derived from Advanced Regional Prediction System (ARPS) weather analyses. The logistic models created from the simulated observations were evaluated using two common statistical measures of model accuracy, the percent correct (PC) and the Hanssen-Kuipers discriminant (HKD). To convert the probabilistic results of the logistic models into a dichotomous yes/no choice suitable for the statistical measures, two critical probability thresholds are considered. The HKD scores are higher when the climatological frequency of contrail occurrence is used as the critical threshold, while the PC scores are higher when the critical probability threshold is 0.5. For both thresholds, typical random errors in temperature, relative humidity, and vertical velocity are found to be small enough to allow for accurate logistic models of contrail occurrence. The accuracy of the models developed from synthetic data is over 85 percent for both the prediction of contrail occurrence and non-occurrence, although in practice, larger errors would be anticipated.

Duda, David P.↗

Data Science and Urban Air Mobility: Challenges and Opportunities

Aviation is broadly a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Urban Air Mobility↗

Data Science and Urban Air Mobility: Challenges and Opportunities

Aviation is broadly a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Urban Air Mobility↗

Data Science Challenges for Urban Air Mobility

Aviation is a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Data Science↗

Data Science Challenges for Urban Air Mobility

Aviation is a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Data Science↗