Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

New-Generation NASA Aura Ozone Monitoring Instrument (OMI) Volcanic SO2 Dataset: Algorithm Description, Initial Results, and Continuation with the Suomi-NPP Ozone Mapping and Profiler Suite (OMPS)

Since the fall of 2004, the Ozone Monitoring Instrument (OMI) has been providing global monitoring of volcanic SO2 emissions, helping to understand their climate impacts and to mitigate aviation hazards. Here we introduce a new-generation OMI volcanic SO2 dataset based on a principal component analysis (PCA) retrieval technique. To reduce retrieval noise and artifacts as seen in the current operational linear fit (LF) algorithm, the new algorithm, OMSO2VOLCANO, uses characteristic features extracted directly from OMI radiances in the spectral fitting, thereby helping to minimize interferences from various geophysical processes (e.g., O3 absorption) and measurement details (e.g., wavelength shift). To solve the problem of low bias for large SO2 total columns in the LF product, the OMSO2VOLCANO algorithm employs a table lookup approach to estimate SO2 Jacobians (i.e., the instrument sensitivity to a perturbation in the SO2 column amount) and iteratively adjusts the spectral fitting window to exclude shorter wavelengths where the SO2 absorption signals are saturated. To first order, the effects of clouds and aerosols are accounted for using a simple Lambertian equivalent reflectivity approach. As with the LF algorithm, OMSO2VOLCANO provides total column retrievals based on a set of predefined SO2 profiles from the lower troposphere to the lower stratosphere, including a new profile peaked at 13 km for plumes in the upper troposphere. Examples given in this study indicate that the new dataset shows significant improvement over the LF product, with at least 50% reduction in retrieval noise over the remote Pacific. For large eruptions such as Kasatochi in 2008 (approximately 1700 kt total SO2/ and Sierra Negra in 2005 (greater than 1100DU maximum SO2), OMSO2VOLCANO generally agrees well with other algorithms that also utilize the full spectral content of satellite measurements, while the LF algorithm tends to underestimate SO2. We also demonstrate that, despite the coarser spatial and spectral resolution of the Suomi National Polar-orbiting Partnership (Suomi-NPP) Ozone Mapping and Profiler Suite (OMPS) instrument, application of the new PCA algorithm to OMPS data produces highly consistent retrievals between OMI and OMPS. The new PCA algorithm is therefore capable of continuing the volcanic SO2 data record well into the future using current and future hyperspectral UV satellite instruments.

OMI↗

NASA Dataset Interoperability Recommendations for Earth Science

NASA ESDS (Earth Science Data and Information System) Dataset Interoperability Working Group has been developing recommendations since 2012 aimed at improving interoperability of EOS (Earth Observing System) datasets. The first set of recommendations were published in 2016 as ESDS RFC-028. The latest set of recommendations is currently undergoing review and will be available as ESDS RFC-036 soon.This talk will inform the ESIP (Earth Science Information Partners) community about the recommendations because their application is relevant to other data producers as well.

Jelenak, Aleksandar↗

Alternative Datasets for Identification of Earth Science Events and Data

Alternative, or non-traditional, data sources can be used to generate datasets which can in turn be analyzed for temporal, spatial and climatological patterns. Events and case studies inferred from the analysis of these patterns can be used by the remote sensing community to more effectively search for Earth observation data. In this paper, we present a new alternative Earth science dataset created from the National Weather Service’s Area Forecast Discussion (AFD) documents. We then present an exploratory methodology for identifying interesting climatological patterns within the AFD data and a corresponding motivating example as to how these data and patterns can be used to search for relevant events or case studies.

Alternative data↗

Alternative Datasets for Identification of Earth Science Events and Data

Alternative, or non-traditional, data sources can be used to generate datasets which can in turn be analyzed for temporal, spatial and climatological patterns. Events and case studies inferred from the analysis of these patterns can be used by the remote sensing community to more effectively search for Earth observation data. In this paper, we present a new alternative Earth science dataset created from the National Weather Service’s Area Forecast Discussion (AFD) documents. We then present an exploratory methodology for identifying interesting climatological patterns within the AFD data and a corresponding motivating example as to how these data and patterns can be used to search for relevant events or case studies.

Bugbee, Kaylin↗

Dataset Documentation FROST: Features Relevant to Ocean Worlds Surface Terrain

We present an analog dataset that provides examples of possible terrain features, geometry, and appearance at the 1-10cm scale on ocean worlds/icy moons such as Europa, Enceladus, and Pluto. The motivation for collecting this dataset was a lack of available high-resolution digital models suitable for development of surface missions to these bodies, including use for simulation of mechanics, sampling, and imaging. NASA field opportunities to Death Valley, California and the Atacama Desert, Chile were leveraged in order to observe and record analog sites.

Wong, Uland↗

Precision Assessment of the HPLC Phytoplankton Pigment Dataset Analyzed by NASA to Quantify Global Variability in Support of Ocean Color Remote Sensing

The ability to generate chlorophyll a (Chl a) assessments from ocean color orbital sensors, such as VIIRS and MODIS, that satisfy the requirements to be climate-quality data record (CDR) quality is contingent in part on the quality of the in situ ground or sea truth observations that serve as datasets for vicarious calibration and algorithm validation activities. NASA has a mandate to collect, analyze, and distribute in situ data of the highest possible quality with documented uncertainties and in keeping with established performance metrics. Using a dataset of over 18,000 HPLC phytoplankton pigment samples representing water collected in all major ocean basins analyzed a central laboratory (Field Support Group (FSG) of the Ocean Ecology Laboratory (OEL) at NASA Goddard Space Flight Center (GSFC)), we performed an assessment of the global precision among sample replicates of Chl a as well as major accessory pigments. We investigated the impacts of filtration volume, water basin, collection technique, pigment concentration, and different filtration volumes for replicate filters on replicate filter precision, as well as investigating any pigment-specific differences. Our results quantify sample variability with the goal of understanding any systemic biases or biogeographic influences.

Thomas, Crystal S.↗

Leveraging Google Earth Engine User Interface for Semiautomated Wetland Classification in the Great Lakes Basin at 10 m With Optical and Radar Geospatial Datasets

As one of the world’s largest freshwater ecosystems,the Great Lakes Basin houses hundreds of thousands of acres of wetlands that support a variety of crucial ecological and environmental functions at the local, regional, and global levels.Monitoring these wetlands is critical to conservation and restoration efforts, however current methods that rely on field monitoring are labor-intensive, costly, and often outdated. In this study, we present a graphical user interface constructed in Google Earth Engine called the Wetland Extent Tool (WET),which allows semi-automatic wetland classification according to a user-input area of interest and date range. WET composites datasets and conducts multi source, moderate resolution processing utilizing Landsat 8 OLI, Sentinel-2 MSI, Sentinel-1 C-SAR, and Shuttle Radar Topography Mission (SRTM) datasets to classify wetlands in the entire Great Lakes Basin. We evaluated classification results of wetlands, uplands, and open water from May-September 2019, and tested whether SRTM elevation, slope,or the Dynamic Surface Water Extent produced the most accurate results in each Great Lake Basin in conjunction with optical indices and radar composites. We found that elevation produced the most accurate classification in Lake Erie, Michigan,and Ontario, while slope performed best in Lake Huron and Superior. Lake Erie, Michigan, Ontario, and Huron achieved high overall accuracy and identification of wetlands. WET leverages cloud-computing for multi source processing of moderate resolution remote sensing data, and employs a user interface in Google Earth Engine that wetland managers and conservationists can use to monitor wetland extent in the Great Lakes Basin in near real-time.

Vanessa L Valenti↗

Biological Insights at the Interface of Multiple Arabidopsis Legacy Datasets

The NASA GeneLab database includes an open-access collection of datasets yielded by space biology experiments. Six Gene Lab Data Sets (GLDS’s) performed in Arabidopsis were selected for analysis (7/17/44/121/205/213), all of which included transcriptome data from spaceflight and ground control environments. Hardware, ecotype, environmental conditions, and other experimental conditions varied, allowing the observations of overarching gene expression impacts of microgravity on plant life without focusing on effects of specific experimental conditions. Using GeneLab pre-processed datasets as the basis for the study, RNA microarray data were analyzed to identify genes that showed altered expression in microgravity when compared to control samples for each individual GLDS. All differentially expressed genes were compared to locate differentially expressed genes common between spaceflight experiments. The most noteworthy result is that not one gene shared differential expression among the six GLDS’s. However, gene expression was not influenced randomly by the microgravity environment, as there were several gene ontology terms that were significantly enriched across all experiments. These included 20 significantly enriched biological processes, and although the genes which enriched each term varied, there were many cases of specific genes common to clusters of multiple GLDS’s. Gene expression such as NAC92 and ERF011 or membrane structural element FFP6 provide insight and direction toward understanding the plant response to spaceflight. Characterizing these common processes and the shared differentially expressed genes has demonstrated potential targets for further study to understand and modulate the biological response of plants in microgravity. Life on Earth has never been subjected to the absence of gravity as a selective pressure, so observing how life forms react to a microgravity environment could provide insight to our shared fundamental biological processes. It is also feasible that the genetic modification of specific genes linked to the microgravity response could improve health and yield of space crops.

Joseph Emhof↗

Automated classification of scientific publications linked to GES DISC datasets

The data collections archived and distributedby the GES DISC NASA data center arewidely utilized for various Earth Science studies.As these collections are created, many researchworks are published regarding the collections, algorithms,validations and applications. SinceGES DISC collects these publications and providestheir citations for the users, it is helpful tocategorize them based on how they relate to the datasetsthey are associated with. Specifically,whether the publication that is linked to GES DISCdataset is using it for applicational research,or if it describes the algorithm for dataset creation,or the validation of the dataset, or providesthe general overview of the data collection. Currently,this process requires simple manuallabelling, and as such, may be possible to solve viaautomation. To approach this problem, wedeveloped machine learning classifiers to predictthe category a publication belongs to. We usedmanually labeled publications as training data forsupervised machine learning algorithms:Random Forest and Naive Bayes. We achieved classificationaccuracy that is substantially betterthan the baseline accuracy, thus greatly improvingthe efficiency of the publication internalanalysis.

Rohan Dayal↗

AgMIP Regional Integrated Assessment of Agricultural Systems in Nioro, Senegal: Representative Agricultural Pathways, Climate, Crop and Economic Datasets

This paper describes the datasets that were used to implement an AgMIP Regional Integrated Assessment for the Nioro region of Senegal to assess the potential impacts of climate change on the principal agricultural system in the Senegal peanut basin and to assess adaptations of that system to climate change under current as well as future climate and socio-economic conditions. This dataset includes the Representative Agricultural Pathways developed for Nioro from 2000-2050; the climate data that were used to implement crop yield simulations; the data that were used to parameterize the DSSAT and APSIM crop models, including historical climate data and future climate scenarios; and the data that were used to parameterize the Tradeoff Analysis Model for Multi-dimensional Impact Assessment (TOAMD) economic simulation model, as well as simulated model outputs.

AgMIP↗

Development of a Global Reference Surface Reflectance and BRDF Datasets from Geostationary Satellite Observations and AERONET Measurements

Surface reflectances and their dependency on illumination-view geometries (i.e., BRDF) are the foundation of many high-level satellite products for land and water monitoring. Yet it is difficult to evaluate the quality of satellite-based surface reflectances with ground-based measurements due to the spatial scale differences. In order to fill the gap, here we develop a reference dataset of surface reflectance and BRDF at the global AERONET sites with data streams from operational geostationary sensors including Himawari 8/9 AHI, GK-2A AMI, and GOES 16/17 ABI. Taking the top-of-atmosphere (TOA) reflectance and the site measured atmospheric aerosol optical depth (AOD) as the main inputs, we apply the GeoNEX-AC algorithm to performance accurate atmospheric correction and derive 10-minute surface reflectance and daily Ross-Thick-Li-Sparse (RTLS) BRDF parameters at AERONET sites where coincident AOD measurements and TOA observations are available from 2016 (for Himawari) or 2018 (for GOES) onwards. The algorithm ensures that the retrieved surface BRDF parameters, along with the site-measured AOD, allow the atmospheric radiative transfer model, SHARM, accurately simulate the observed TOA reflectance at diurnal and longer time scales. They are our best estimates of the surface optical properties and thus can serve as the “reference” to evaluate the performance of operational atmospheric correction algorithms (where AOD is assumed unknown and needs to be retrieved). The reference BRDF also allow us to evaluate the spectral band ratios between the SWIR (e.g., 2200 nm) and the visible (e.g., 650 nm) regions, which are commonly used in operational atmospheric correction algorithms. Finally, we demonstrate that the reference dataset can be used to develop potential data synergies between different GEO satellites as well as GEO-LEO sensors.

Weile Wang↗

Interpretable Convolutional Learning Classifier System (C-LCS) for Higher Dimensional Datasets

The purpose of this paper is to devise an interpretable hybrid classification model for Convolutional Neural Networks (CNN) and a Learning Classifier System (LCS). The presented hybrid system integrates the fundamental attributes from both types of these classifiers. In the proposed hybrid model CNN works as an automatic feature extractor, and LCS works to provide interpretable rule-based classification results. Although LCS has limitations working on higher dimensional datasets, we resolve this limitation by using CNN as a feature extractor. The other concept of the non-interpretability of CNN is addressed by using the LCS rule. Furthermore, our experiment with higher dimensional datasets like CIFAR-10 and Fashion-MNIST shows that extended LCS provides comparable performance to the standard neural network model while also providing interpretable results. We named this extended LCS method Convolutional Learning Classifier Cystem (C-LCS).

Jelani Owens↗

Quantifying radiation quality for space relevant radiation types: Fitting excess risk models to three combined HZE-irradiated mouse datasets

Radiation health risks are predominantly derived from low linear energy transfer (LET) terrestrial exposures; however, space radiation includes exposure to high-LET and high-charge, high-energy (HZE) particles. Accurately quantifying the differences in radiation quality between the space and terrestrial radiation environments is important for assessing and predicting health risks for astronauts. Weil et al. 2009 and 2014 used two different inbred mouse strains to study differences in hepatocellular carcinoma (HCC) tumorigenesis after exposures to low- and high- LET radiation[1-2]. More recently, Edmondson et al. 2020 provided valuable new tumor data in outbred mice that were exposed to low- and high-LET radiation[3]. The present study aims to rigorously investigate a relative biological effectiveness (RBE) factor by leveraging the HCC tumor data from the three datasets[1-3]. The three experiments were similarly designed, allowing the raw data to be combined into a pooled dataset to estimate excess relative risk (ERR) and excess absolute risk (EAR) models using Bayesian Poisson regression.

Lori J. Chappell↗

Development of the Ames Global Hyperspectral Synthetic Dataset

This study develops the surface BRDF (bidirectional reflectance distribution function) product of the Ames Global Hyperspectral Synthetic Dataset (AGHSD), based on the corresponding MODIS products, to support the NASA Surface Biology and Geology mission development. A main challenge in deriving a hyperspectral dataset from the multi-band satellite products is how to identify a succinct yet robust algorithm that allow us to infer BRDF at unobserved wavelengths based on the few observed bands. Using the theories of radiative transfer in vegetation canopies, we arrive at a simple equation that accurately approximates hyperspectral surface BRDF as the weighted sum of components from the soil and the vegetation. Each of the components is modeled by the product of the spectrally-dependent optical properties of a surface element (the spectra of the soil surface reflectance, the leaf single albedo, or the canopy scattering coefficient) and a spectrally-independent bidirectional scattering function. The optical properties of the soil and the vegetation can be obtained from existing spectral libraries or model simulations. The bidirectional scattering functions are represented by the Ross-Thick-Li-Sparse BRDF model, where the linear coefficients are estimated with regression analysis from the multi-band MODIS data. We validate the algorithm with simulations by Monte Carlo Ray Tracing model experiments, and the results are highly consistent with the theoretic derivation. We apply the algorithm to generate the AGHSD BRDF product at 1km and 8-day resolutions for the year of 2019. The results are biogeochemically and physically coherent and consistent, and thus serve the goal to support the science and application development of the SBG community.

Hyperspectral↗

Newly Released GPCP Version 3.2 Global Precipitation Datasets at NASA GES DISC

The Global Precipitation Climatology Project (GPCP) is the precipitation component of an internationally coordinated set of (mainly) satellite-based global products dealing with the Earth's water and energy cycles, under the auspices of the Global Water and Energy Experiment (GEWEX) Data and Assessment Panel (GDAP) of the World Climate Research Program. As the follow-on to the GPCP Version 2.X products, GPCP Version 3 (GPCP V3.2) seeks to continue the production of long, homogeneous precipitation record using modern input and calibration datasets. The GPCP V3.2 provides globally complete analyses of surface precipitation on a 0.5°x 0.5° latitude/longitude grid at both monthly and daily intervals, respectively covering 1983 to the present and June 2000 to the present. New data fields have been introduced to better characterize the precipitation, particularly including an estimate of the fraction of the precipitation that is liquid (rain) in both the Monthly and Daily, and a Quality Index for the Monthly. Compared to the operational GPCP V2.3 Monthly, the V3.2 Monthly provides a more reasonable climatology in the Southern Ocean, and increases the global average precipitation by about 4.46%, which is in line with recommendations of recent assessments. However, the two versions have comparable global and regional trends for 1983-2020. Compared to the operational One-Degree Daily (Version 1.3) product, the V3.2 Daily better represents the histogram of precipitation rates, particularly at high values. In this presentation, we will present the latest GPCP V3.2 daily and monthly datasets archived and distributed at NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC) along with examples from NASA Giovanni.

Precipitation↗

Spatially-Coordinated Airborne Data and Complementary Products for Aerosol, Gas, Cloud, and Meteorological Studies: the Nasa Activate Dataset

The NASA Aerosol Cloud meTeorology Interactions oVer the western ATlantic Experiment (ACTIVATE) produced a unique dataset for research into aerosol–cloud–meteorology interactions, with applications extending from process-based studies to multi-scale model intercomparison and improvement as well as to remote-sensing algorithm assessments and advancements. ACTIVATE used two NASA Langley Research Center aircraft, a HU-25 Falcon and King Air, to conduct systematic and spatially coordinated flights over the northwest Atlantic Ocean, resulting in 162 joint flights and 17 other single-aircraft flights between 2020 and 2022 across all seasons. Data cover 574 and 592 cumulative flights hours for the HU-25 Falcon and King Air, respectively. The HU-25 Falcon conducted profiling at different level legs below, in, and just above boundary layer clouds (< 3 km) and obtained in situ measurements of trace gases, aerosol particles, clouds, and atmospheric state parameters. Under cloud-free conditions, the HU-25 Falcon similarly conducted profiling at different level legs within and immediately above the boundary layer. The King Air (the high-flying aircraft) flew at approximately ∼ 9 km and conducted remote sensing with a lidar and polarimeter while also launching dropsondes (785 in total). Collectively, simultaneous data from both aircraft help to characterize the same vertical column of the atmosphere. In addition to individual instrument files, data from the HU-25 Falcon aircraft are combined into “merge files” on the publicly available data archive that are created at different time resolutions of interest (e.g., 1, 5, 10, 15, 30, 60 s, or matching an individual data product's start and stop times). This paper describes the ACTIVATE flight strategy, instrument and complementary dataset products, data access and usage details, and data application notes. The data are publicly accessible through https://doi.org/10.5067/SUBORBITAL/ACTIVATE/DATA001 (ACTIVATE Science Team, 2020).

Aerosol Cloud meTeorology Interactions oVer the we↗

GLORIA - A Globally Representative Hyperspectral In Situ Dataset for Optical Sensing of Water Quality

The development of algorithms for remote sensing of water quality (RSWQ) requires a large amount of in situ data to account for the bio-geo-optical diversity of inland and coastal waters. The GLObal Reflectance community dataset for Imaging and optical sensing of Aquatic environments (GLORIA) includes 7,572 curated hyperspectral remote sensing reflectance measurements at 1 nm intervals within the 350 to 900 nm wavelength range. In addition, at least one co-located water quality measurement of chlorophyll a , total suspended solids, absorption by dissolved substances, and Secchi depth, is provided. The data were contributed by researchers affiliated with 59 institutions worldwide and come from 450 different water bodies, making GLORIA the de-facto state of knowledge of in situ coastal and inland aquatic optical diversity. Each measurement is documented with comprehensive methodological details, allowing users to evaluate fitness-for-purpose, and providing a reference for practitioners planning similar measurements. We provide open and free access to this dataset with the goal of enabling scientific and technological advancement towards operational regional and global RSWQ monitoring.

remote sensing of water quality↗

A Machine Learning Ready Dataset of Acoustic Power Maps for Detection of Active Region Emergence

The development of an accurate forecast for solar eruptive activity has become increasingly important in order to prevent any potential impact on activities in space and the Earth's environment. It is therefore crucial to detect active regions before they appear on the solar surface and create early warning capabilities for upcoming Space Weather disturbances. In this work, 9TB of solar data (SDO/HMI dopplergrams, magnetograms and continuum intensity maps) involving the emergence of 61 NOAA solar active regions since 2010 were processed using the NASA HECC capabilities. An acoustic power maps time-series dataset was created (for four different frequency ranges and processed to take into account the solar sphere geometric effect ) which can be used for understanding the dynamics of the solar surface and train a variety of ML models. The calculated acoustic power maps carry precursor information associated with the decrease in continuum intensity on the solar surface, verifying older helioseismology research. Our results show that a Long Short-Term Memory (LSTMs) model, with a modest layer depth and the right hyperparameters tuned, when trained on this solar acoustic power maps dataset can predict without false negatives a drop in intensity (associated with the emergence of the active region), up to 18 hours in advance.

SMD↗