Search NASA⌕ Search

SEARCH · Search NASA

Results for “synthetic dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Development of the Ames Global Hyperspectral Synthetic Dataset

This study develops the surface BRDF (bidirectional reflectance distribution function) product of the Ames Global Hyperspectral Synthetic Dataset (AGHSD), based on the corresponding MODIS products, to support the NASA Surface Biology and Geology mission development. A main challenge in deriving a hyperspectral dataset from the multi-band satellite products is how to identify a succinct yet robust algorithm that allow us to infer BRDF at unobserved wavelengths based on the few observed bands. Using the theories of radiative transfer in vegetation canopies, we arrive at a simple equation that accurately approximates hyperspectral surface BRDF as the weighted sum of components from the soil and the vegetation. Each of the components is modeled by the product of the spectrally-dependent optical properties of a surface element (the spectra of the soil surface reflectance, the leaf single albedo, or the canopy scattering coefficient) and a spectrally-independent bidirectional scattering function. The optical properties of the soil and the vegetation can be obtained from existing spectral libraries or model simulations. The bidirectional scattering functions are represented by the Ross-Thick-Li-Sparse BRDF model, where the linear coefficients are estimated with regression analysis from the multi-band MODIS data. We validate the algorithm with simulations by Monte Carlo Ray Tracing model experiments, and the results are highly consistent with the theoretic derivation. We apply the algorithm to generate the AGHSD BRDF product at 1km and 8-day resolutions for the year of 2019. The results are biogeochemically and physically coherent and consistent, and thus serve the goal to support the science and application development of the SBG community.

Hyperspectral↗

Neural Network Atmospheric Correction of Remote Sensing Imagery Over Water Using a Synthetic Dataset

Remote sensing atmospheric correction methods have primarily focused on imagery over land. However, accurate correction over water is important for monitoring and research of aquatic environments. More research in this area is ongoing, though one of the biggest challenges is enough quality data to develop and validate correction methods. This is especially true for neural network (NN) -based models which have shown promise in this area given enough quality data. To address this deficiency of data, we are leveraging a synthetic dataset produced by a model called SWIPE that uses radiative transfer modeling to simulate the atmospheric effects on water-leaving (WL) reflectance to estimate top-of-atmosphere (TOA) reflectance. This allows us to produce almost unlimited pairs of WL reflectance and corresponding TOA reflectance for model training across a variety of atmospheric conditions. We use two approaches for our atmospheric correction model. One uses a conditional variational autoencoder (VAE) to estimate a single WL reflectance value from a single TOA reflectance value. The second is based on a UNET architecture and estimates an array of WL reflectance values from an array of TOA reflectance values. The goal of the second method is to capture atmospheric effects that occur spatially between values within the array as compared to the first method.

deep learning↗

Synthetic Hyperspectral Data for Global Water Quality Algorithm Development

Eutrophication and increasing prevalence of potentially toxic algal blooms (cyanoHABs) among global inland water bodies have become a major ecological concern and require direct attention. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local aquatic processes, to spatially resolved global products. Planned aquatic biogeochemistry remote sensing data products from hyperspectral imagers such as NASA’s Surface Biology and Geology (SBG) mission and relevant aquatic sensor sensitivity precursor airborne imaging spectrometer data provide unprecedented radiometric resolution and sensor sensitivity for characterizing complex aquatic ecosystems. However, scarcity of high-quality freshwater in-situ optical data hinders our capability to develop and validate robust retrieval algorithms. A state-of-the-art synthetic dataset of paired top-of-atmosphere, bottom-of-atmosphere, and optical and biogeophysical data was developed through radiative transfer modeling to simulate natural freshwater ecosystems. A synthetic or precursor dataset for SBG is being used to train robust machine learning models to derive water quality products pertinent to SBG mission objectives. The dataset is also used to show the potential of performing vigorous aquatic sensitivity studies and explored pathways for how best to optimize hyperspectral data for machine learning development. A processing pipeline and resultant global synthetic/precursor dataset for inland waters is presented to establish the innovation for water quality studies of inland waters globally. Optical Society of America Imaging and Applied Optics Congress, Hyperspectral Imaging and Sounding of the Environment (OSA HISE) Meeting, 19-23 July 2021, Virtual Meeting, https://www.osa.org/enus/meetings/osa_meetings/optical_sensors_and_sensing_congress/program/hyperspectral_imaging_and_sounding_of_the_environm/

Synthetic↗

Aircraft Engine Run-To-Failure Data Set Under Real Flight Conditions

The generation of data-driven prognostics models requires the availability of datasets with run-to-failure trajectories. In order to contribute to the development of these methods, the dataset provides a new realistic dataset of run-to-failure trajectories for a small fleet of aircraft engines under realistic flight conditions. The damage propagation modelling used for the generation of this synthetic dataset builds on the modelling strategy from previous work [1] and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet [2]. Secondly, it extends the degradation modelling by relating the degradation process to the operation history. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) dynamical model [3]. More details about the generation process can be found in [4].

CMAPSS↗

LISA Framework for Enhancing Gravitational Wave Signal Extraction Techniques

This paper describes the development of a Framework for benchmarking and comparing signal-extraction and noise-interference-removal methods that are applicable to interferometric Gravitational Wave detector systems. The primary use is towards comparing signal and noise extraction techniques at LISA frequencies from multiple (possibly confused) ,gravitational wave sources. The Framework includes extensive hybrid learning/classification algorithms, as well as post-processing regularization methods, and is based on a unique plug-and-play (component) architecture. Published methods for signal extraction and interference removal at LISA Frequencies are being encoded, as well as multiple source noise models, so that the stiffness of GW Sensitivity Space can be explored under each combination of methods. Furthermore, synthetic datasets and source models can be created and imported into the Framework, and specific degraded numerical experiments can be run to test the flexibility of the analysis methods. The Framework also supports use of full current LISA Testbeds, Synthetic data systems, and Simulators already in existence through plug-ins and wrappers, thus preserving those legacy codes and systems in tact. Because of the component-based architecture, all selected procedures can be registered or de-registered at run-time, and are completely reusable, reconfigurable, and modular.

Thompson, David E.↗

Dimensionality Reduction Through Classifier Ensembles

In data mining, one often needs to analyze datasets with a very large number of attributes. Performing machine learning directly on such data sets is often impractical because of extensive run times, excessive complexity of the fitted model (often leading to overfitting), and the well-known "curse of dimensionality." In practice, to avoid such problems, feature selection and/or extraction are often used to reduce data dimensionality prior to the learning step. However, existing feature selection/extraction algorithms either evaluate features by their effectiveness across the entire data set or simply disregard class information altogether (e.g., principal component analysis). Furthermore, feature extraction algorithms such as principal components analysis create new features that are often meaningless to human users. In this article, we present input decimation, a method that provides "feature subsets" that are selected for their ability to discriminate among the classes. These features are subsequently used in ensembles of classifiers, yielding results superior to single classifiers, ensembles that use the full set of features, and ensembles based on principal component analysis on both real and synthetic datasets.

Oza, Nikunj C.↗

Bayesian Retrieval of Complete Posterior PDFs of Oceanic Rain Rate From Microwave Observations

This paper presents a new Bayesian algorithm for retrieving surface rain rate from Tropical Rainfall Measurements Mission (TRMM) Microwave Imager (TMI) over the ocean, along with validations against estimates from the TRMM Precipitation Radar (PR). The Bayesian approach offers a rigorous basis for optimally combining multichannel observations with prior knowledge. While other rain rate algorithms have been published that are based at least partly on Bayesian reasoning, this is believed to be the first self-contained algorithm that fully exploits Bayes Theorem to yield not just a single rain rate, but rather a continuous posterior probability distribution of rain rate. To advance our understanding of theoretical benefits of the Bayesian approach, we have conducted sensitivity analyses based on two synthetic datasets for which the true conditional and prior distribution are known. Results demonstrate that even when the prior and conditional likelihoods are specified perfectly, biased retrievals may occur at high rain rates. This bias is not the result of a defect of the Bayesian formalism but rather represents the expected outcome when the physical constraint imposed by the radiometric observations is weak, due to saturation effects. It is also suggested that the choice of the estimators and the prior information are both crucial to the retrieval. In addition, the performance of our Bayesian algorithm is found to be comparable to that of other benchmark algorithms in real-world applications, while having the additional advantage of providing a complete continuous posterior probability distribution of surface rain rate.

Chiu, J. Christine↗

Simulating Visible/Infrared Imager Radiometer Suite Normalized Difference Vegetation Index Data Using Hyperion and MODIS

The success of MODIS (the Moderate Resolution Imaging Spectrometer) in creating unprecedented, timely, high-quality data for vegetation and other studies has created great anticipation for data from VIIRS (the Visible/Infrared Imager Radiometer Suite). VIIRS will be carried onboard the joint NASA/Department of Defense/National Oceanic and Atmospheric Administration NPP (NPOESS (National Polar-orbiting Operational Environmental Satellite System) Preparatory Project). Because the VIIRS instruments will have lower spatial resolution than the current MODIS instruments 400 m versus 250 m at nadir for the channels used to generate Normalized Difference Vegetation Index data, scientists need the answer to this question: how will the change in resolution affect vegetation studies? By using simulated VIIRS measurements, this question may be answered before the VIIRS instruments are deployed in space. Using simulated VIIRS products, the U.S. Department of Agriculture and other operational agencies can then modify their decision support systems appropriately in preparation for receipt of actual VIIRS data. VIIRS simulations and validations will be based on the ART (Application Research Toolbox), an integrated set of algorithms and models developed in MATLAB(Registerd TradeMark) that enables users to perform a suite of simulations and statistical trade studies on remote sensing systems. Specifically, the ART provides the capability to generate simulated multispectral image products, at various scales, from high spatial hyperspectral and/or multispectral image products. The ART uses acquired ( real ) or synthetic datasets, along with sensor specifications, to create simulated datasets. For existing multispectral sensor systems, the simulated data products are used for comparison, verification, and validation of the simulated system s actual products. VIIRS simulations will be performed using Hyperion and MODIS datasets. The hyperspectral and hyperspatial properties of Hyperion data will be used to produce simulated MODIS and VIIRS products. Hyperion-derived MODIS data will be compared with near-coincident MODIS collects to validate both spectral and spatial synthesis, which will ascertain the accuracy of converting from MODIS to VIIRS. MODIS-derived VIIRS data is needed for global coverage and for the generation of time series for regional and global investigations. These types of simulations will have errors associated with aliasing for some scene types. This study will help quantify these errors and will identify cases where high-quality, MODIS-derived VIIRS data will be available.

Ross, Kenton W.↗

Land ice height-retrieval algorithm for NASA's ICESat-2 photon-counting laser altimeter

The Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) and its sole scientific instrument, the Advanced Topographic Laser Altimeter System (ATLAS), was launched on 15 September 2018 with a primary goal of measuring changes in the surface of the Earth's land ice (glaciers and ice sheets). ATLAS is a photon-counting laser altimeter, which records the transit time of individual photons in order to reconstruct surface height along track. The ground-track pattern repeats every 91 days such that changes in ice sheet surface height can be estimated through time. In this paper, we describe the set of algorithms that have been developed for ICESat-2 to retrieve ice sheet surface height from the geolocated photons for the Land Ice Along-Track Height Product (ATL06), and demonstrate their output and performance using a synthetic dataset over various land-ice surfaces and under different cloud conditions. We show that the ATL06 algorithm is expected to perform at the level required to meet the ICESat-2 science objectives for land ice.

ICESat-2↗

Remotely Estimating Total Suspended Solids Concentration in Clear to Extremely Turbid Waters Using a Novel Semi-Analytical Method

Total suspended solids (TSS) concentration is an important biogeochemical parameter for water quality management and sediment-transport studies. In this study, we propose a novel semi-analytical method for estimating TSS in clear to extremely turbid waters from remote-sensing reflectance (Rrs). The proposed method includes three sub-algorithms used sequentially. First, the remotely sensed waters are classified into clear (Type I), moderately turbid (Type II), highly turbid (Type III), and extremely turbid (Type IV) water types by comparing the values of Rrs at 490, 560, 620, and 754 nm. Second, semi-analytical models specific to each water type are used to determine the particulate backscattering coefficients (bbp) at a corresponding single wavelength (i.e., 560 nm for Type I, 665 nm for Type II, 754 nm for Type III, and 865 nm for Type IV). Third, a specific relationship between TSS and bbp at the corresponding wavelength is used in each water type. Unlike other existing approaches, this method is strictly semi-analytical and its sub-algorithms were developed using synthetic datasets only. The performance of the proposed method was compared to that of three other state-of-the-art methods using simulated (N = 1000, TSS ranging from 0.01 to 1100 g/m3) and in situ measured (N = 3421, TSS ranging from 0.09 to 2627 g/m3) pairs of Rrs and TSS. Results showed a significant improvement with a Median Absolute Percentage Error (MAPE) of 16.0% versus 30.2–90.3% for simulated data and 39.7% versus 45.9–58.1% for in situ data, respectively. The new method was subsequently applied to 175 MEdium Resolution Imaging Spectrometer (MERIS) and 498 Ocean and Land Colour Instrument (OLCI) images acquired in the 2003–2020 timeframe to produce long-term TSS time-series for Lake Suwa and Lake Kasumigaura, Japan. Performance assessments using MERIS and OLCI matchups showed good agreements with in situ TSS measurements.

Mulit-Wavelength↗

Instantaneous Photosynthetically Available Radiation (IPAR) prediction models based on Neural Network for ocean waters.

Instantaneous photosynthetically available radiation (IPAR) at the ocean surface and its vertical profile below the surface play a critical role in models to calculate net primary productivity of marine phytoplankton. In this work, we report two IPAR prediction models based on neural network (NN) approach, one for open ocean and the other for coastal waters. These models are trained, validated, and tested using a large volume of synthetic datasets for open ocean and coastal waters simulated by a radiative transfer model. Our NN models are designed to predict IPAR under a large range of atmospheric and oceanic conditions. The NN models can compute subsurface IPAR profile very accurately up to euphotic zone depth. The root mean square errors associated with the diffuse attenuation coefficient of IPAR are less than 0.011 𝑚−1 and 0.036 𝑚−1 for open ocean and coastal waters respectively. The performance of the NN models is better than presently available semi analytical models, with significant superiority in coastal waters.

PACE↗

Improving Sim-to-Real Transfer in Vision-Based Robot Navigation Via Instance-Level GAN-Based Data Augmentation

Achieving robust vision-based robotic tasks requires large amounts of data, which are often difficult to obtain in real-world scenarios. Simulators and synthetic data offer a cost-effective alternative, but the visual gap between simulation and reality hinders the performance of models when deployed in real-world environments. In this paper, we present a data augmentation pipeline that integrates a foundation model (Segment Anything Model) with an unsupervised image-to-image translation model (CycleGAN) for instance-level domain transfer from simulation to reality. This pipeline enables the generation of realistic labeled data from synthetic images for training supervised machine learning models in vision-based navigation tasks. We evaluate our approach on real-world data for ego-vehicle pose estimation, a critical autonomous navigation task involving the prediction of cross-track position and heading angle relative to road center line markings. The results of our tests show that our GAN-based data augmentation pipeline significantly outperforms models trained solely on simulation data or on data processed with standard image augmentation methods for sim-to-real transfer, enhancing model robustness and generalizability in real-world scenarios. Our method provides a scalable and flexible data augmentation tool for leveraging large synthetic datasets to enhance vision-based robotic navigation tasks.

artificial intelligence↗

Optimizing a Small RNAseq Analysis Pipeline for NASA GeneLab Using Open-Source Tools and Libraries

Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.

SmallRNAseq, NASA GeneLab, quality control, adapte↗

Earth Science Imagery Registration

The study of global environmental changes involves the comparison, fusion, and integration of multiple types of remotely-sensed data at various temporal, radiometric, and spatial resolutions. Results of this integration may be utilized for global change analysis, as well as for the validation of new instruments or for new data analysis. Furthermore, future multiple satellite missions will include many different sensors carried on separate platforms, and the amount of remote sensing data to be combined is increasing tremendously. For all of these applications, the first required step is fast and automatic image registration, and as this need for automating registration techniques is being recognized, it becomes necessary to survey all the registration methods which may be applicable to Earth and space science problems and to evaluate their performances on a large variety of existing remote sensing data as well as on simulated data of soon-to-be-flown instruments. In this paper we present one of the first steps toward such an exhaustive quantitative evaluation. First, the different components of image registration algorithms are reviewed, and different choices for each of these components are described. Then, the results of the evaluation of the corresponding algorithms combining these components are presented o n several datasets. The algorithms are based on gray levels or wavelet features and compute rigid transformations (including scale, rotation, and shifts). Test datasets include synthetic data as well as data acquired over several EOS Land Validation Core Sites with the IKONOS and the Landsat-7 sensors.

LeMoigne, Jacqueline↗

Simulation Results of the Huygens Probe Entry and Descent Trajectory Reconstruction Algorithm

Cassini/Huygens is a joint NASA/ESA mission to explore the Saturnian system. The ESA Huygens probe is scheduled to be released from the Cassini spacecraft on December 25, 2004, enter the atmosphere of Titan in January, 2005, and descend to Titan s surface using a sequence of different parachutes. To correctly interpret and correlate results from the probe science experiments and to provide a reference set of data for "ground-truthing" Orbiter remote sensing measurements, it is essential that the probe entry and descent trajectory reconstruction be performed as early as possible in the postflight data analysis phase. The Huygens Descent Trajectory Working Group (DTWG), a subgroup of the Huygens Science Working Team (HSWT), is responsible for developing a methodology and performing the entry and descent trajectory reconstruction. This paper provides an outline of the trajectory reconstruction methodology, preliminary probe trajectory retrieval test results using a simulated synthetic Huygens dataset developed by the Huygens Project Scientist Team at ESA/ESTEC, and a discussion of strategies for recovery from possible instrument failure.

Kazeminejad, B.↗

The Mock LISA Data Challenges: History, Status, Prospects

This slide presentation reviews the importance for the Mock LISA Data Challenges (MLDC). Laser Interferometer Space Antenna (LISA) is a gravitational wave (GW) observatory that will return data such that data analysis is integral to the measurement concept. Further rationale of the MLDC are to kickstart the development of a LISA data-analysis computational infrastructure, and to encourage, track, and compare progress in LISA data-analysis development in the open community. The MLDCs is a coordinated, voluntary effort in GW community, that will periodically issue datasets with synthetic noise and GW signals from sources of undisclosed parameters; increasing difficulty. The challenge participants return parameter estimates and descriptions of search methods. Some of the challenges and the resultant entries are reviewed. The aim is to show that LISA data analysis is possible, and to develop new techniques, using multiple international teams for the development of LISA core analysis tools

data analysis↗

On the Feasibility of Monitoring Carbon Monoxide in the Lower Troposphere from a Constellation of Northern Hemisphere Geostationary Satellites (PART 1)

By the end of the current decade, there are plans to deploy several geostationary Earth orbit (GEO) satellite missions for atmospheric composition over North America, East Asia and Europe with additional missions proposed. Together, these present the possibility of a constellation of geostationary platforms to achieve continuous time-resolved high-density observations over continental domains for mapping pollutant sources and variability at diurnal and local scales. In this paper, we use a novel approach to sample a very high global resolution model (GEOS-5 at 7 km horizontal resolution) to produce a dataset of synthetic carbon monoxide pollution observations representative of those potentially obtainable from a GEO satellite constellation with predicted measurement sensitivities based on current remote sensing capabilities. Part 1 of this study focuses on the production of simulated synthetic measurements for air quality OSSEs (Observing System Simulation Experiments). We simulate carbon monoxide nadir retrievals using a technique that provides realistic measurements with very low computational cost. We discuss the sampling methodology: the projection of footprints and areas of regard for geostationary geometries over each of the North America, East Asia and Europe regions; the regression method to simulate measurement sensitivity; and the measurement error simulation. A detailed analysis of the simulated observation sensitivity is performed, and limitations of the method are discussed. We also describe impacts from clouds, showing that the efficiency of an instrument making atmospheric composition measurements on a geostationary platform is dependent on the dominant weather regime over a given region and the pixel size resolution. These results demonstrate the viability of the "instrument simulator" step for an OSSE to assess the performance of a constellation of geostationary satellites for air quality measurements.

GEOS-5↗

Application of Support Vector Regression to Derive Crater Depth/Diameter From Satellite Images

Through the study of impact crater shapes, one can draw important conclusions about the nature and evolution of planetary surfaces [e.g., 1-4].In particular, studying the depth (d) to diameter (D)ratio (d/D) of a population of impact craters, in combination with crater count statistics, can yield valuable insights regarding rates of erosion and burial[5]. Motivated by the great abundance of available planetary surface image data, the goal of this project is to develop an efficient way to estimate d/D from satellite images of impact craters for which stereo information is not available [6]. We set out to develop and train a machine learning algorithm to extract d/D from a dataset of synthetic impact crater images for which model d/D is known. The applications of machine learning to planetary science are numerous and diverse [7], including automatic planetary surface mapping [8] and the detection of impact craters [9]. Our algorithm makes use of Support Vector Regression (SVR), which is a type of Support Vector Machine (SVM) [10, 11].SVMs are a branch of supervised machine learning valued for their straightforward implementation and versatility in solving both classification and regression problems. In regression analysis, an SVR algorithm produces a hyperplane function to fit the training data points, as well as an ε-tube that surrounds the hyperplane. Tunable hyperparameters include the width of the ε-tube (ε) and the amount an algorithm is penalized for points which fall outside the ε-tube.

L R Chin↗