Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Commercial Synthetic Aperture Radar Data for Surface Deformation and Change

The commercial synthetic aperture radar (SAR) market is experiencing year-over-year growth and is currently capable of imaging anywhere in the world within an hour at X-band. Surface Deformation and Change (SDC) is a mission study at NASA to investigate innovative architectures beyond the NASA-ISRO SAR (NISAR) mission in the next decade. In this article we present preliminary findings on the technical capabilities of the currently available commercial SAR data, and investigate their applicability to SDC goals, such as imaging quality and retrieval of surface motion.

SAR↗

DESI DR1 Ly α 1D power spectrum: the Fast Fourier Transform estimator measurement

Here, we present the one-dimensional Lyman-α forest power spectrum measurement derived from the data release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI). The measurement of the Lyman-α forest power spectrum along the line of sight from high-redshift quasar spectra provides information on the shape of the linear matter power spectrum, neutrino masses, and the properties of dark matter. In this work, we use a Fast Fourier Transform (FFT)-based estimator, which is validated on synthetic data in a companion paper. Compared to the FFT measurement performed on the DESI early data release, we improve the noise characterization with a cross-exposure estimator and test the robustness of our measurement using various data splits. We also refine the estimation of the uncertainties and now present an estimator for the covariance matrix of the measurement. Furthermore, we compare our results to previous high-resolution and eBOSS measurements. In another companion paper, we present the same DR1 measurement using the Quadratic Maximum Likelihood Estimator (QMLE). These two measurements are consistent with each other and constitute the most precise one-dimensional power spectrum measurement to date, while being in good agreement with results from the DESI early data release.

Lyman alpha forest↗

An Integrated Centroid Finding and Particle Overlap Decomposition Algorithm for Stereo Imaging Velocimetry

An integrated algorithm for decomposing overlapping particle images (multi-particle objects) along with determining each object s constituent particle centroid(s) has been developed using image analysis techniques. The centroid finding algorithm uses a modified eight-direction search method for finding the perimeter of any enclosed object. The centroid is calculated using the intensity-weighted center of mass of the object. The overlap decomposition algorithm further analyzes the object data and breaks it down into its constituent particle centroid(s). This is accomplished with an artificial neural network, feature based technique and provides an efficient way of decomposing overlapping particles. Combining the centroid finding and overlap decomposition routines into a single algorithm allows us to accurately predict the error associated with finding the centroid(s) of particles in our experiments. This algorithm has been tested using real, simulated, and synthetic data and the results are presented and discussed.

McDowell, Mark↗

FracML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage

Poster on “FRACML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. The accurate characterization of subsurface fracture networks is essential for the secure operation of carbon capture, utilization, and storage (CCUS) projects. A thorough understanding of the spatial distribution of subsurface faults and fractures is crucial for predicting CO2 plume evolution and minimizing risks such as potential leakage into overlying formations or induced seismicity. In this context, robust fracture network quantification plays a pivotal role in reservoir management, providing the data necessary to fine-tune operational parameters, and ensure the environmental and economic viability of CCUS projects. As part of the U.S. Department of Energy’s SMART (Science-informed Machine Learning for Accelerating Real-time Decisions in Subsurface Applications) initiative, we focused on the development and application of a machine learning-based tool (FRACML) designed to quantify and map fracture networks using real-world (non-synthetic) data from an active CO2 injection site. Our objective is to demonstrate the utility of this tool in improving operational efficiency and safety across CCUS sites.

artifical intelligence / machine learning (AI/ML)↗

Artificial Intelligence Medical Support for Long-Duration Space Missions

We envision an artificial intelligence (AI) based system that will provide support and recommendations to the crew medical officer (CMO) and ground flight surgeon during long-duration space missions. Such a system would be pretrained on the knowledgebase of clinical knowledge on Earth, minimizing the amount of Earth data that needs to be transferred into space. Then during deployment, the system would be constantly refined through active learning from diverse streams of data from sensors in the spacecraft, data collected daily from individual astronauts, and human-in-the-loop feedback from the crew. The model could be interrogated for predictions and recommendations on personalized crew health based on the overall status of the spacecraft, medicinal stores, and status of other crew members. Adaptation techniques would be used to incorporate spaceflight data that have very different distributions from the training data due to the extreme environment. Edge computing and the most advanced neuromorphic processing would enable computation in scenarios with low power and bandwidth, while dimensionality reduction would be employed to ensure that the input data streams from spaceflight are as small as possible. In order to realize this long-term vision, several hardware and software aspects need to be developed and assembled. First, models pretrained on Earth biomedical data would need to be evaluated for predictive accuracy, and the best one selected. That model would need to be adapted to learn from diverse, sparse, and inconsistently measured data streams, as well as human-in-the-loop feedback. A data integration, standardization, and dimensionality reduction methodology would need to be developed to handle all data types and feed them into the model. Once the software and data infrastructure is developed, it would need to be integrated with small footprint compute processors and tested in high-radiation, high-vibration, unregulated temperature situations. As a short-term goal, we recommend to focus on the development of the data and model software structure. Several large language models (LLM) already exist that have been trained on Earth biomedical and clinical knowledgebases, including BioMedLLM, Med-PaLM, SPOKE LLM, and Foresight. These models need to be evaluated for accuracy and the best one chosen for a proof-of-concept structure, while maintaining awareness of the accelerating AI field and incorporating any newly improved model architectures as needed. Then, we recommend to develop a database of synthetic data types to mimic the diverse data streams that are expected in a long-duration space mission. This should include environmental and microbial data from the spacecraft, non-invasive data from wearables and point-of-care devices employed by astronauts, and more invasive molecular and physiological monitoring of clinical and biomarker data from astronauts. The data standardization methodology should be developed, and these data streams used to refine the clinical LLM. Several scenarios should be developed that could plausibly come up in a long-duration space mission, and changes or aberrations introduced to the data at specific times to mimic these scenarios. Then, question and answer tasks should be designed to interrogate the model for predictions and recommendations, with acceptable answers already identified.

Artificial Intelligence↗

Evaluating quenching in cosmological simulations of galaxy formation with spectral covariance in the optical window

ABSTRACT Cosmological hydrodynamical simulations provide valuable insights on galaxy evolution when coupled with observational data. Comparisons with real galaxies are typically performed via scaling relations of the observables. Here, we follow an alternative approach based on the spectral covariance in a model-independent way. We build upon previous work by Sharbaf et al. that studied the covariance of high-quality SDSS (Sloan Digital Sky Survey) continuum-subtracted spectra in a relatively narrow range of velocity dispersion ($\sigma \in [100,150]$ km s$^{-1}$). Here, the same analysis is applied to synthetic data from the eagle and IllustrisTNG100 simulations, to assess the ability of these runs to mimic real galaxies. The real and simulated spectra are consistent regarding spectral covariance, although with subtle differences that can inform the implementation of subgrid physics. Spectral fitting done a posteriori on stacks segregated with respect to latent space reveals that the first principal component (PC1) is predominantly influenced by the stellar age distribution, with an underlying age–metallicity degeneracy. Good agreement is found regarding star formation prescriptions but there is disagreement with active galactic nucleus (AGN) feedback, that also affects the subset of quiescent galaxies. We show a substantial difference in the implementation of the AGN subgrid prescriptions, regarding central black hole seeding, that could lead to the mismatch. Differences are manifest between these two simulations in the star formation histories stacked with respect to latent space. We emphasize that this methodology only relies on the spectral covariance to assess whether simulations provide a true representation of galaxy formation.

Sharbaf, Z. (ORCID:0009000450545946)↗

Studies in astronomical time series analysis. IV - Modeling chaotic and random processes with linear filters

While chaos arises only in nonlinear systems, standard linear time series models are nevertheless useful for analyzing data from chaotic processes. This paper introduces such a model, the chaotic moving average. This time-domain model is based on the theorem that any chaotic process can be represented as the convolution of a linear filter with an uncorrelated process called the chaotic innovation. A technique, minimum phase-volume deconvolution, is introduced to estimate the filter and innovation. The algorithm measures the quality of a model using the volume covered by the phase-portrait of the innovation process. Experiments on synthetic data demonstrate that the algorithm accurately recovers the parameters of simple chaotic processes. Though tailored for chaos, the algorithm can detect both chaos and randomness, distinguish them from each other, and separate them if both are present. It can also recover nonminimum-delay pulse shapes in non-Gaussian processes, both random and chaotic.

Scargle, Jeffrey D.↗

Earth Science Imagery Registration

The study of global environmental changes involves the comparison, fusion, and integration of multiple types of remotely-sensed data at various temporal, radiometric, and spatial resolutions. Results of this integration may be utilized for global change analysis, as well as for the validation of new instruments or for new data analysis. Furthermore, future multiple satellite missions will include many different sensors carried on separate platforms, and the amount of remote sensing data to be combined is increasing tremendously. For all of these applications, the first required step is fast and automatic image registration, and as this need for automating registration techniques is being recognized, it becomes necessary to survey all the registration methods which may be applicable to Earth and space science problems and to evaluate their performances on a large variety of existing remote sensing data as well as on simulated data of soon-to-be-flown instruments. In this paper we present one of the first steps toward such an exhaustive quantitative evaluation. First, the different components of image registration algorithms are reviewed, and different choices for each of these components are described. Then, the results of the evaluation of the corresponding algorithms combining these components are presented o n several datasets. The algorithms are based on gray levels or wavelet features and compute rigid transformations (including scale, rotation, and shifts). Test datasets include synthetic data as well as data acquired over several EOS Land Validation Core Sites with the IKONOS and the Landsat-7 sensors.

LeMoigne, Jacqueline↗

Can Selforganizing Maps Accurately Predict Photometric Redshifts?

We present an unsupervised machine-learning approach that can be employed for estimating photometric redshifts. The proposed method is based on a vector quantization called the self-organizing-map (SOM) approach. A variety of photometrically derived input values were utilized from the Sloan Digital Sky Survey's main galaxy sample, luminous red galaxy, and quasar samples, along with the PHAT0 data set from the Photo-z Accuracy Testing project. Regression results obtained with this new approach were evaluated in terms of root-mean-square error (RMSE) to estimate the accuracy of the photometric redshift estimates. The results demonstrate competitive RMSE and outlier percentages when compared with several other popular approaches, such as artificial neural networks and Gaussian process regression. SOM RMSE results (using delta(z) = z(sub phot) - z(sub spec)) are 0.023 for the main galaxy sample, 0.027 for the luminous red galaxy sample, 0.418 for quasars, and 0.022 for PHAT0 synthetic data. The results demonstrate that there are nonunique solutions for estimating SOM RMSEs. Further research is needed in order to find more robust estimation techniques using SOMs, but the results herein are a positive indication of their capabilities when compared with other well-known methods

Way, Michael J.↗

Radiative Studies of Planetary Atmospheres

Retrieval algorithms and associated software for application to CIRS infrared spectral data have been developed and coded. A general forward radiative transfer code has been written that runs efficiently on a Macintosh, even at high spectral resolution (0.5 per centimeter). It makes use of the correlated-k approach for representation of the gaseous absorption and can include those gases listed in the HITRAN and GEISA atlases, along with collision-induced absorption. Cloud effects are included as spectrally dependent absorbers. Provision has been made for future extension to include particle scattering in an n-stream approximation. The primary purpose of the code is to produce synthetic data and to serve as the forward calculating element in gas and cloud retrieval programs developed for the Mac as well as other platforms. Initial development of algorithms and production software suitable for application to CIRS data to be obtained from Jupiter, Saturn and Titan has been completed, and production versions of the software for application to the spectral data are in place. This includes temperature, gaseous constituent, and cloud opacity retrieval, algorithms that can be applied to both nadir and limb data. This work has been done as a cooperative effort between Conrath and Matcheva (Cornell), Achterberg (GSFWSSAI), and Flasar (GSFC).

Conrath, Barney J.↗

Using a Genetic Algorithm to Model Broadband Regional Waveforms for Crustal Structure in the Western United States

In this study, we analyze regional seismograms to obtain the crustal structure in the eastern Great Basin and western Colorado plateau. Adopting a for- ward-modeling approach, we develop a genetic algorithm (GA) based parameter search technique to constrain the one-dimensional crustal structure in these regions. The data are broadband three-component seismograms recorded at the 1994-95 IRIS PASSCAL Colorado Plateau to Great Basin experiment (CPGB) stations and supplemented by data from U.S. National Seismic Network (USNSN) stations in Utah and Nevada. We use the southwestern Wyoming mine collapse event (M(sub b) = 5.2) that occurred on 3 February 1995 as the seismic source. We model the regional seismograms using a four-layer crustal model with constant layer parameters. Timing of teleseismic receiver functions at CPGB stations are added as an additional constraint in the modeling. GA allows us to efficiently search the model space. A carefully chosen fitness function and a windowing scheme are added to the algorithm to prevent search stagnation. The technique is tested with synthetic data, both with and without random Gaussian noise added to it. Several separate model searches are carried out to estimate the variability of the model parameters. The average Colorado plateau crustal structure is characterized by a 40-km-thick crust with velocity increases at depths of about 10 and 25 km and a fast lower crust while the Great Basin has approximately 35- km-thick crust and a 2.9-km-thick sedimentary layer.

Bhattacharyya, Joydeep↗

Multilabel proportion prediction and out-of-distribution detection on gamma spectra of short-lived fission products

In the machine learning problem of multilabel classification, the objective is to determine for each test instance which classes the instance belongs to. In this work, we consider an extension of multilabel classification, called multilabel proportion prediction, in the context of radioisotope identification (RIID) using gamma spectra data. We aim to not only predict radioisotope proportions, but also identify out-of-distribution (OOD) spectra. We achieve this goal by viewing gamma spectra as discrete probability distributions, and based on this perspective, we develop a custom semi-supervised loss function that combines a traditional supervised loss with an unsupervised reconstruction error function. Our approach was motivated by its application to the analysis of short-lived fission products from spent nuclear fuel. In particular, we demonstrate that a neural network model trained with our loss function can successfully predict the relative proportions of 37 radioisotopes simultaneously. The model trained with synthetic data was then applied to measurements taken by Pacific Northwest National Laboratory (PNNL) to conduct analysis typically done by subject-matter experts. Here, we also extend our approach to successfully identify when measurements are OOD, and thus should not be trusted, whether due to the presence of a novel source or novel proportions.

Anomaly detection↗

Digital Twin + AI: Control Room of the Future [Slides]

The control room functions as the central brain of the grid, essential for balancing supply and demand and ensuring moment-to-moment grid reliability. Like the human brain, which processes sensory data to make decisions, control room operators analyze operational data from power generation, transmission, and distribution to make informed decisions. Currently, decision-making primarily rests with operators due to hardware and software limitations. However, with technological advancements, Digital Twins and AI are becoming high interest points in the control room's decision-making pilot programs. NREL is developing a comprehensive decision-making platform that integrates Digital Twins, AI, and advanced visualization techniques. As this integration progresses, the role of Digital Twins will evolve from conducting automated simulations to serving as a Trustworthy AI enabler, offering verification and validation of AI-generated response for power systems or providing physics-aware synthetic data of AI pre-training.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A three-point velocity estimation method for two-dimensional coarse-grained imaging data

Time delay and velocity estimation methods have been widely studied subjects in the context of signal processing, with applications in many different fields of physics. The velocity of waves or coherent fluctuation structures is commonly estimated as the distance between two measurement points divided by the time lag that maximizes the cross correlation function between the measured signals, but this is demonstrated to result in erroneous estimates for two spatial dimensions. We present an improved method to accurately estimate both components of the velocity vector, relying on three non-aligned measurement points. We introduce a stochastic process describing the fluctuations as a superposition of uncorrelated pulses moving in two dimensions. Using this model, we show that the three-point velocity estimation method, using time delays calculated through cross correlations, yields the exact velocity components when all pulses have the same velocity. The two- and three-point methods are tested on synthetic data generated from realizations of such processes for which the underlying velocity components are known. The results reveal the superiority of the three-point technique. Finally, we demonstrate the applicability of the velocity estimation on gas puff imaging data of strongly intermittent plasma fluctuations due to the radial motion of coherent, blob-like structures at the boundary of the Alcator C-Mod tokamak.

Materials Science↗

A Fast Implementation of the ISODATA Clustering Algorithm

Clustering is central to many image processing and remote sensing applications. ISODATA is one of the most popular and widely used clustering methods in geoscience applications, but it can run slowly, particularly with large data sets. We present a more efficient approach to ISODATA clustering, which achieves better running times by storing the points in a kd-tree and through a modification of the way in which the algorithm estimates the dispersion of each cluster. We also present an approximate version of the algorithm which allows the user to further improve the running time, at the expense of lower fidelity in computing the nearest cluster center to each point. We provide both theoretical and empirical justification that our modified approach produces clusterings that are very similar to those produced by the standard ISODATA approach. We also provide empirical studies on both synthetic data and remotely sensed Landsat and MODIS images that show that our approach has significantly lower running times.

Memarsadeghi, Nargess↗

A Fast Implementation of the Isodata Clustering Algorithm

Clustering is central to many image processing and remote sensing applications. ISODATA is one of the most popular and widely used clustering methods in geoscience applications, but it can run slowly, particularly with large data sets. We present a more efficient approach to IsoDATA clustering, which achieves better running times by storing the points in a kd-tree and through a modification of the way in which the algorithm estimates the dispersion of each cluster. We also present an approximate version of the algorithm which allows the user to further improve the running time, at the expense of lower fidelity in computing the nearest cluster center to each point. We provide both theoretical and empirical justification that our modified approach produces clusterings that are very similar to those produced by the standard ISODATA approach. We also provide empirical studies on both synthetic data and remotely sensed Landsat and MODIS images that show that our approach has significantly lower running times.

Memarsadeghi, Nargess↗

Total Variation Majorization Minimization (TV-MM) Approach to Radiometer Brightness Temperature Gridding and Reconstruction

This paper presents the implementation of an algorithm to enhance the image resolution of the Earth's surface brightness temperature (T B ) data measured by radiometers such as the one onboard of the Soil Moisture Active Passive (SMAP) mission. A key step in radiometer T B processing is the conversion of the swath-based calibrated antenna temperature (T A ) measurements to the Level 3 Earth-centered grid. The simplest algorithm to transform this data from swath to gridded format is called drop-in-the-bucket which simply averages surrounding noisy T A samples to form a T B value at the gridded location. This method reduces noise, however produces low resolution products. To obtain a higher resolution product, SMAP uses other techniques such the Backus-Gilbert (BG) algorithm, which is the conventional method used in microwave radiometry. Although this method performs the required interpolation, it is not effective in denoising and removing blurring effects due to antenna filtering of the radiometer image data. Our motivation for this development is to further improve the resolution through post-processing of the radiometer T B image, a highly cost-effective method of image enhancement. The approach adapted in this work is based on the minimization of the Total Variation (TV) regularized objective function that is used extensively in solving general ill-posed linear inverse problems in image processing. Since the TV-based objective function is convex but not everywhere differentiable, there exists many numerical algorithms that can estimate the solution and the one selected for this work is called Majorization- Minimization (MM). By applying this algorithm, simulation experiments were performed based on synthetic data from the Geophysical model as well as real SMAP data to demonstrate the effectiveness of the technique. Results were then compared against the BG method.

Wing Lee↗

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.

97 MATHEMATICS AND COMPUTING↗