Search NASA⌕ Search

SEARCH · Search NASA

Results for “dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Microstructure Segmentation with Deep Learning Encoders Pre-Trained on a Large Microscopy Dataset

This study examined the improvement of microscopy segmentation accuracy by transfer learning from a large dataset of microscopy images called MicroNet. Many neural network encoder architectures, including VGG, Inception, and ResNet, were trained on over 100,000 labelled microscopy images from 54 classes. These pre-trained encoders were then embedded into multiple segmentation architectures including U-Net and DeepLabV3+ to evaluate segmentation performance on newly created benchmark microscopy datasets. Compared to ImageNet pre-training, models pre-trained on MicroNet generalized better to out-of-distribution micrographs taken under different imaging and sample conditions and were more accurate with less training data. When training with only a single Ni-superalloy image, pre-training on MicroNet produced a 72.2 percent reduction in relative segmentation error. These results suggest that transfer learning from large in-domain datasets generate models with learned feature representations that are more useful for downstream tasks and will likely improve any microscopy image analysis technique that can leverage pre-trained encoders.

machine learning↗

Scaling photosynthetic function and CO2 dynamics from leaf to canopy level for maize – dataset combining diurnal and seasonal measurements of vegetation fluorescence, reflectance and vegetation indices with canopy gross ecosystem productivity

Recent advances in leaf fluorescence measurements and canopy proximal remote sensing currently enable the non-destructive collection of rich diurnal and seasonal time series, which are required for monitoring vegetation function at the temporal and spatial scales relevant to the natural dynamics of photosynthesis. Remote sensing assessments of vegetation function have traditionally used actively excited foliar chlorophyll fluorescence measurements, canopy optical reflectance data and vegetation indices (VIs), and only recently passive solar induced chlorophyll fluorescence (SIF) measurements. In general, reflectance data are more sensitive to the seasonal variations in canopy chlorophyll content and foliar biomass, while fluorescence observations more closely relate to the dynamic changes in plant photosynthetic function. With this dataset we link leaf level actively excited chlorophyll fluorescence, canopy proximal reflectance and SIF, with eddy covariance measurements of gross ecosystem productivity (GEP). The dataset was collected during the 2017 growing season on maize, using three automated systems (i.e., Monitoring Pulse-Amplitude-Modulation fluorimeter, Moni-PAM; Fluorescence Box, FloX; and from eddy covariance tower). The data were quality checked, filtered and collated to a common 30 minutes timestep. We derived vegetation indices related to canopy functioning (e.g., Photochemical Reflectance Index, PRI; Normalized Difference Vegetation Index, NDVI; Chlorophyll Red-edge, Clre) to investigate how SIF and VIs can be coupled for monitoring vegetation photosynthesis. The raw datasets and the filtered and collated data are provided to enable new processing and analyses.

Petya Campbell↗

AI4MARS: A Dataset for Terrain-Aware Autonomy on Mars

Deep learning has quickly become a necessity for selfdriving vehicles on Earth. In contrast, the self-driving vehicles on Mars, including NASA’s latest rover, Perseverance, which is planned to land on Mars in February 2021, are still driven by classical machine vision systems. Deep learning capabilities, such as semantic segmentation and object recognition, would substantially benefit the safety and productivity of ongoing and future missions to the red planet. To this end, we created the first large-scale dataset, AI4Mars, for training and validating terrain classification models for Mars, consisting of ~326K semantic segmentation full image labels on 35K images from Curiosity, Opportunity, and Spirit rovers, collected through crowdsourcing. Each image was labeled by ~10 people to ensure greater quality and agreement of the crowdsourced labels. It also includes ~1.5K validation labels annotated by the rover planners and scientists from NASA’s MSL (Mars Science Laboratory) mission, which operates the Curiosity rover, and MER (Mars Exploration Rovers) mission, which operated the Spirit and Opportunity rovers. We trained a DeepLabv3 model on the AI4Mars training dataset and achieved over 96% overall classification accuracy on the test set. The dataset is made publicly available.1

Ono, Hiro↗

Microstructure Segmentation With Deep Learning Encoders Pre-Trained on a Large Microscopy Dataset

This study examined the improvement of microscopy segmentation intersection over union accuracy by transfer learning from a large dataset of microscopy images called MicroNet. Many neural network encoder architectures were trained on over 100,000 labeled microscopy images from 54 material classes. These pre-trained encoders were then embedded into multiple segmentation architectures including UNet and DeepLabV3+ to evaluate segmentation performance on created benchmark microscopy datasets. Compared to ImageNet pre-training, models pre-trained on MicroNet generalized better to out-of-distribution micrographs taken under different imaging and sample conditions and were more accurate with less training data. When training with only a single Ni-superalloy image, pre-training on MicroNet produced a 72.2% reduction in relative intersection over union error. These results suggest that transfer learning from large in-domain datasets generate models with learned feature representations that are more useful for downstream tasks and will likely improve any microscopy image analysis technique that can leverage pre-trained encoders.

machine learning↗

Making Dataset Quality Information FAIR: Supporting Open-Source Science and Enhancing (Re)Use and Trustworthiness of Scientific Data

- Quality information should be documented and readily shared within and across domains. - Sharing of dataset quality information supports open science and trustworthiness of scientific data. - Dataset quality is more than data quality. - Quality tends to be domain-specific and context-dependent. - Community guidelines provide practical steps towards FAIR dataset quality information.

Ge Peng↗

Recommendations on Funding Mission Operations and Historical Datasets

The Heliophysics Low Cost Access to Space (H-LCAS) and Flight Opportunities in Research and Technology (H-FORT) grant budgets primarily fund the development and construction of instrumentation. A side effect of this approach is that mission operations, including data collection and data processing, tend to be severely under budget (or unfunded). In order to satisfy the requirements of modern missions, we recommend a new funding source for mission operations. This funding source is increasingly vital going forward as new missions collect exponentially more data compared to past missions. Similarly, to take full advantage of underutilized historical datasets, we recommend adding another funding source to analyze these valuable datasets. Without these funding sources, mission datasets will, in the best case, be significantly underdeveloped and underutilized, and more likely, will fall dramatically short of their required scope.

Mykhaylo Shumko↗

The Multiplatform Precipitation Feature (MPF) Database: Synthesizing Satellite and Ground-Based Precipitation and Lightning Datasets for Convective Studies

NASA’s Lightning Imaging Sensor (LIS) and the Global Precipitation Measurement (GPM) mission have contributed a wealth of data toward global lightning and precipitation studies, respectively. Combining lightning and precipitation datasets leverages their unique insights into deep convective processes that inform about characteristics of convection and its intensity. Recent efforts to synthesize the LIS and GPM datasets prepare the opportunity for unprecedented large-scale, value-added multiplatform analyses of convection. This data synthesis proof-of-concept study elaborates on the creation of a database of reflectivity-based multiplatform precipitation features (MPFs) that capture a combination of information extracted from spatiotemporally coincident lightning and precipitation data within individual storm features. The space-based GPM Dual-frequency Precipitation Radar (DPR) provides a record of precipitation data, while the GPM Validation Network (VN) additionally incorporates ground-based polarimetric Doppler radar data to provide microphysical and kinematic context to DPR data. The LIS instrument onboard the International Space Station has contributed lightning observations since 2017. MPFs encapsulating information from these datasets are created from isolated regions of filtered, smoothed DPR reflectivity data to which ellipses are fit. Each MPF includes feature location, size, and eccentricity information as well as summary reflectivity characteristics. They also include summaries of precipitation microphysics and derived three-dimensional wind available from ground-based radar data. LIS data provides standard lightning characteristics such as flash count and density to each MPF as well as other informative metrics such as flash area and radiance. Each MPF file includes information about the original data from which the MPF and its characteristics were determined, allowing end-user reconstruction of the ellipse and deeper “level I” analysis of captured data. This database of VN-LIS MPFs enables broad statistical analysis of the relationships between the microphysical, kinematic, and electrical properties of convection. Preliminary results from a demonstration of the database will be described as well as ongoing efforts and avenues for future work.

Lightning↗

Low Latency Flux and Concentration Datasets in Support of Greenhouse Gas Monitoring Based on NASA's GEOS Modeling and Data Assimilation System

We present efforts to develop space-based greenhouse gas monitoring systems that can provide low latency information and traceability to independent observations. Through support from its Carbon Monitoring System program, NASA has developed the capability to assimilate XCO2 retrievals from the Orbiting Carbon Observatory, 2 (OCO-2) into the Goddard Earth Observing System (GEOS) Constituent Data Assimilation System (CoDAS) to create gap-filled, three-dimensional (3D) estimates of CO2 mixing ratio. When OCO-2 data are not available, concentration fields are further informed by a bottom-up flux package based on remotely sensed fire radiative power, nighttime lights, and vegetation reflectance combined with estimates of atmospheric growth rate based on surface in situ data. The 3D nature of this dataset supports evaluation with independent aircraft data, helping to ensure transparency of remotely sensed data products. These quasi-operational data are currently produced 2-3 months behind real time and are distributed via NASA and international dashboard services to a variety of end users. In this presentation, we provide an overview of the system as well as remaining data gaps and modeling challenges. We also highlight the application of this dataset for detecting emissions anomalies associated with COVID-19 and comparing against independent emissions estimates. Finally, we highlight a new NASA initiative called the Earth Information System (EIS), which aims to support open science and applications by leveraging emerging cloud computing capabilities to increase access to NASA’s greenhouse gas datasets, opportunities for co-development, and transparency in methods for analysis and flux attribution.

Lesley Ott↗

Curating AI-Ready Datasets for Equity and Environmental Justice: A Data-Centric AI Case Study

An equitable and environmentally just community is essentialin order to avoid disproportionate burden borne by vulnerablecommunities. This need becomes pressing in the aftermathof an extreme event such as disaster or hazard when it is diffi-cult for the governing bodies to implement resource allocationas per the need. Artificial Intelligence (AI) algorithms canhelp surface Equity and Environmental Justice (EEJ) issueswhen trained on EEJ datasets. However, curating AI-readyEEJ training datasets is challenging due to differences in fac-tors such as heterogeneity, resolution, modality, and level ofexpertise in labeling. Additionally, EEJ issues involve sensi-tive information where uncertainties and errors could degradethe performance of AI algorithms. For eg. Error in seasonalcrop yield information can highly affect the prediction of an-nual crop yield. To address these challenges, Data-centricAI (DCAI) methods are employed, which enhance AI algo-rithm performance even with limited training samples. DCAIprioritizes data quality, thereby reducing the adverse effectsof uncertainties and errors during the model training process.This research proposes a novel dataset and benchmark for an-alyzing the effect of the Maui Wildfire of 2023 for Equityand Environmental Justice (EEJ) issues. The proposed datasetaligns with the concepts of DCAI such as annotation quality,data preprocessing, privacy, feature engineering, governanceand provenance. We firmly believe that the proposed datasetwould lay a foundation to implement robust and reliable mod-ern AI algorithms for addressing EEJ issues.

Paridhi Parajuli↗

NETL RDE Image Classification Dataset 2025 - 14 Classes

Dataset including high-speed down-axis RDE images used for updated image classification study. This dataset includes 180,000 images with 14 classifications: 1CW, 1CCW, 2CW, 2CCW, 3CW, 3CCW, Deflagration, 4CW, 4CCW, 5CW, and 5CCW. Images are cropped to center annulus, and resized to 301x301 pixels. Images are filtered using the AFRL Beta correction factor.

Dataset↗

NETL RDE Image Classification Dataset 2020 - 10 Classes

Dataset including high-speed down-axis RDE images used for updated image classification study. This dataset includes 100,000 images with 10 classifications: 1CW, 1CCW, 2CW, 2CCW, 3CW, 3CCW, and Deflagration. Images are cropped to center annulus, and resized to 301x301 pixels. Images are filtered using the AFRL Beta correction factor.

AS↗

FlowDash Geothermal Energy Enhancer: Where is Next Geothermal Resource? Machine Learning + Multiple Datasets => Geothermal Exploration Indication?

This is the presentation delivered at the 2025 GEODE Datathon competition. GEODE is a consortium of experts that addresses technology and knowledge gaps in geothermal energy, leveraging technology and best practices from the oil and gas industry. NETL team was awarded the 1st place in the engineering track. 2025 GEODE Datathon had a total of 42 teams from top universities and several major industrial companies. This awarded work is founded on a robust idea and innovative approach that uses machine learning coupled to multiple datasets to visualize geothermal “sweet” spots/indications in Great Basin based on the data provided from the GEODE Datathon. The use case also leveraged other datasets and demonstrated insightful and valuable indications for geothermal exploration.

Geothermal energy, Machine learning, Multiple Data↗

A Forward-Looking Dataset of EV Managed Charging Resource and Costs

This presentation summarizes a high-resolution, forward-looking dataset of EV adoption, EV charging, and managed charging resource. Vehicle-level data are grounded in current adoption and charging patterns, and ~200,000 real-world vehicle-weeks of travel data covering all on-road segments (i.e., light-duty, transit and school buses, local, regional and long-haul medium- and heavy-duty). The data, which include multiple charging profiles per vehicle to bound flexibility, are then processed and aggregated to describe baseline charging and charge management resource by county, hour, year, scenario, and vehicle type. Coupled with one of four scenarios of how EV managed charging costs might evolve over time, the dataset enables a power sector capacity expansion model to select cost-optimal quantities of EV managed charging and supply-side resources to reliably satisfy demand. Five integration strategies: Baseline, Daytime and Flat (passive), Flex (active), and Stress (anti-strategy), illustrate how baseline charging and flexibility potential changes with EVSE build-out and charging preferences.

33 ADVANCED PROPULSION SYSTEMS↗

The Foundational Industry Energy Dataset: Unit-level Characterization and Derived Energy Estimates for Industrial Facilities in 2017

The Foundational Industry Energy Dataset (FIED) addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization by facility. Each facility is identified by a unique registryID, based on the U.S. Environmental Protection Agency (EPA) Facility Registry Service, and includes its coordinates and other geographic identifiers. Energy-using units are characterized by design capacity, as well as their estimated energy use, greenhouse gas emissions, and physical throughput using 2017 data from the EPA's National Emissions Inventory and Greenhouse Gas Reporting Program. An overview of the derivation methods is provided in a separate technical report which will be linked after publication. The Python code used to compile the dataset is available in a GitHub repository. An updated 2020 version is under development.

Array↗

Historical Bolide Infrasound Dataset (1960–1972)

We present the first fully curated, publicly accessible archive of infrasonic records from ten large bolide events documented by the U.S. Air Force Technical Applications Center’s global microbarometer network between 1960 and 1972. Captured on analog strip-chart paper, these waveforms predate modern digital arrays and space-based sensors, making them a unique window on meteoroid activity in the mid-twentieth century. Prior studies drew important scientific conclusions from the records but released only limited artifacts, chiefly period–amplitude tables and unprocessed scans, leaving the underlying data inaccessible for independent study. The present release transforms those limited excerpts into a research-ready resource. By capturing ten large events in the mid-20th century, the dataset constitutes a critical reference point for assessing bolide activity before the advent of modern space-based and digital ground-based monitoring. The multi-year coverage and worldwide distribution of events provide a valuable reference for comparing past and more recent detections, facilitating assessments of long-term flux and the dynamics of acoustic wave propagation in Earth’s atmosphere. The dataset’s availability in a consolidated format ensures straightforward access to waveforms and derived measurements, supporting a wide range of scientific inquiries into bolide physics and infrasound monitoring. By preserving these historical acoustic observations, the collection maintains a significant record of mid-20th-century meteoroid entries. It thereby establishes a basis for further refinement of impact hazard evaluations, contributes to historical continuity in atmospheric observation, and enriches the study of meteoroid-generated infrasound signals on a global scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Comparison of GFED3, QFED2 and FEER1 Biomass Burning Emissions Datasets in a Global Model

Biomass burning contributes about 40% of the global loading of carbonaceous aerosols, significantly affecting air quality and the climate system by modulating solar radiation and cloud properties. However, fire emissions are poorly constrained in models on global and regional levels. In this study, we investigate 3 global biomass burning emission datasets in NASA GEOS5, namely: (1) GFEDv3.1 (Global Fire Emissions Database version 3.1); (2) QFEDv2.4 (Quick Fire Emissions Dataset version 2.4); (3) FEERv1 (Fire Energetics and Emissions Research version 1.0). The simulated aerosol optical depth (AOD), absorption AOD (AAOD), angstrom exponent and surface concentrations of aerosol plumes dominated by fire emissions are evaluated and compared to MODIS, OMI, AERONET, and IMPROVE data over different regions. In general, the spatial patterns of biomass burning emissions from these inventories are similar, although the strength of the emissions can be noticeably different. The emissions estimates from QFED are generally larger than those of FEER, which are in turn larger than those of GFED. AOD simulated with all these 3 databases are lower than the corresponding observations in Southern Africa and South America, two of the major biomass burning regions in the world.

biomass burning↗

Outcomes of a NASA Workshop to Develop a Portfolio of Low Latency Datasets for Time-Sensitive Applications

It is widely accepted that time-sensitive remote sensing data serve the needs of decision makers in the applications communities and yet to date, a comprehensive portfolio of NASA low latency datasets has not been available. This paper will describe the NASA low latency, or Near-Real Time (NRT), portfolio, how it was developed and plans to make it available online through a portal that leverages the existing EOSDIS capabilities such as the Earthdata Search Client (https:search.earthdata.nasa.gov), the Common Metadata Repository (CMR) and the Global Imagery Browse Service (GIBS). This paper will report on the outcomes of a NASA Workshop to Develop a Portfolio of Low Latency Datasets for Time-Sensitive Applications (27-29 September 2016 at NASA Langley Research Center, Hampton VA). The paper will also summarize findings and recommendations from the meeting outlining perceived shortfalls and opportunities for low latency research and application science.

remote sensing↗