Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks associated with deep space crewed missions (cis-Lunar, Mars transit/surface) require development of health countermeasures, novel ecosystem support, risk modeling, and fundamental space biological knowledge discovery. Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from space biological and health studies are needed for reuse by scientists to address these tasks. The data as well as space-relevant biospecimens are being made more findable, accessible, interoperable, and reusable through NASA’s Open Science Data Repository (OSDR). This new OSDR umbrella grouping includes NASA GeneLab, the NASA Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection. The OSDR system design appropriately handles metadata and processed-tabular results from ALSDA studies collected from space experiments. But raw and processed ALSDA bioimage and video datasets require an expansion of OSDR’s data architecture to handle ingestion, curation, and egress. The academic-industry bioimaging field saw a scientific renaissance in the past several years through leveraging open-source software, international collaborations, machine learning, and other open science/programming approaches. As crewed missions and more biological experiments are on the deep space horizon, OSDR is embracing data stewardship through listening to feedback from subject matter experts and designing an expanded architecture which is appropriate for NASA’s goals to enable analysis and reuse of bioimaging and video data for the public science community.Discovery Through Image and Video Data Sharing

space biology↗

Diesel Fuel Consumption in Prominent U.S. Open-Pit Mines: Site-Level Estimates

This report presents a comprehensive framework for estimating diesel fuel consumption and prices at open-pit mines in the United States. The framework includes transparent methods for calculating site-level diesel energy use when direct reporting is unavailable, and a structured confidence evaluation for each method. The framework is demonstrated to estimate current diesel consumption at 21 open-pit mines in the United States. Initial findings support ongoing efforts to strengthen the competitiveness and security of the U.S. industrial base by supporting data-driven supply chain analysis and decision-making, improved transparency in mining sector energy use, and targeted deployment of energy innovation and cost-reduction strategies. Future updates to the dataset—coupled with expanded data transparency and method validation—will help ensure that the findings remain relevant as the sector continues to evolve.

02 PETROLEUM↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

TRMM Precipitation Application Examples Using Data Services at NASA GES DISC

Data services to support precipitation applications are important for maximizing the NASA TRMM (Tropical Rainfall Measuring Mission) and the future GPM (Global Precipitation Mission) mission's societal benefits. TRMM Application examples using data services at the NASA GES DISC, including samples from users around the world will be presented in this poster. Precipitation applications often require near-real-time support. The GES DISC provides such support through: 1) Providing near-real-time precipitation products through TOVAS; 2) Maps of current conditions for monitoring precipitation and its anomaly around the world; 3) A user friendly tool (TOVAS) to analyze and visualize near-real-time and historical precipitation products; and 4) The GES DISC Hurricane Portal that provides near-real-time monitoring services for the Atlantic basin. Since the launch of TRMM, the GES DISC has developed data services to support precipitation applications around the world. In addition to the near-real-time services, other services include: 1) User friendly TRMM Online Visualization and Analysis System (TOVAS; URL: http://disc2.nascom.nasa.gov/Giovanni/tovas/); 2) Mirador (http://mirador.gsfc.nasa.gov/), a simplified interface for searching, browsing, and ordering Earth science data at GES DISC. Mirador is designed to be fast and easy to learn; 3) Data via OPeNDAP (http://disc.sci.gsfc.nasa.gov/services/opendap/). The OPeNDAP provides remote access to individual variables within datasets in a form usable by many tools, such as IDV, McIDAS-V, Panoply, Ferret and GrADS; and 4) The Open Geospatial Consortium (OGC) Web Map Service (WMS) (http://disc.sci.gsfc.nasa.gov/services/wxs_ogc.shtml). The WMS is an interface that allows the use of data and enables clients to build customized maps with data coming from a different network.

Liu, Zhong↗

Noble Gases and Stable Isotopes Track the Origin and Early Evolution of the Venus Atmosphere

The composition the atmosphere of Venus results from the integration of many processes entering into play over the entire geological history of the planet. Determining the elemental abundances and isotopic ratios of noble gases (He, Ne, Ar, Kr, Xe) and stable isotopes (H, C, N, O, S) in the Venus atmosphere is a high priority scientific target since it could open a window on the origin and early evolution of the entire planet. This chapter provides an overview of the existing dataset on noble gases and stable isotopes in the Venus atmosphere. The current state of knowledge on the origin and early and long-term evolution of the Venus atmosphere deduced from this dataset is summarized. A list of persistent and new unsolved scientific questions stemming from recent studies of planetary atmospheres (Venus, Earth and Mars) are described. Important mission requirements pertaining to the measurement of volatile elements in the atmosphere of Venus as well as potential technical difficulties are outlined.

noble gases↗

Fostering Open Science in Earth Data Science Research: Insights From Earthdata Forum By ASDC

In the dynamic landscape of Earth Science research, the promotion of open science principles is paramount for advancing knowledge and collaboration. The Earthdata Forum is an actively maintained and operational user forum for all participating National Aeronautics and Space Administration (NASA) Earth Observing System Data and Information System (EOSDIS) Distributed Active Archive Centers (DAACs), and the Global Change Master Directory (GCMD). The Forum serves as a cross-DAAC platform from which user communities can obtain authoritative information relating to NASA Earth Science. This abstract explores the role of the Earthdata Forum forum.earthdata.nasa.gov as a pivotal platform in fostering open science within the Earth Science community. The platform serves as a hub for researchers to actively engage in discussions, share datasets, and collaboratively tackle challenges in the field. Key aspects discussed include the platform's contribution to data accessibility, collaboration, and knowledge sharing. Forum.earthdata.nasa.gov provides a space where researchers transparently ask questions, discuss methodologies, share insights, and seek advice from a vibrant community. The resulting collaborative environment not only facilitates the exchange of ideas but also bolsters the collective knowledge base.

Earthdata FORUM↗

NASA Environmental Justice Data Search Interface Overview

NASA’s Earth Science Division (ESD) is committed to empower Environmental Justice (EJ) communities by expanding awareness, accessibility, and use of Earth science data to enable contributions to Earth science research and applications. To that end, the NASA Earth Science Data Systems (ESDS) Program developed an EJ Data Catalog, a simple guide to NASA datasets and socioeconomic datasets that may be useful in EJ research. The EJ Data Catalog is divided by topics—such as disasters, urban flooding, extreme heat, food availability, water availability, climate, and health and air quality—and possible use cases for each dataset. The new version of the EJ Data Catalog is now integrated into NASA’s Science Discovery Engine (SDE), an open-source science infrastructure to enable collaborative and interdisciplinary science. In this workshop you will learn about NASA’s Equity and Environmental Justice (EEJ) activities and opportunities as well as participate on an interactive live demo of the new Science Discovery Engine for Environmental Justice.

environmental justice↗

The NASA Open Science Data Repository: Biomedical Data, Analysis Tools, and Informatic Collaborations

Increased biomedical risks and challenges associated with deep space missions require knowledge discovery, health countermeasures, and biomedical support capabilities. Maximally open-access and reusable data is needed by developers, scientists, and engineers to develop these systems. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database (ie., findable, accessible, interoperable, and reusable), and meets various scientific, technical, and operational needs. It offers users and submitters the ability to upload, download, search, share, analyze, cite, and visualize data across ‘omics, physiological, phenotypic, payload, hardware, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR is an expanded database, based upon the successes of NASA GeneLab. OSDR has >460 studies with datasets covering model organisms to non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets with raw files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) developed from industry norms. OSDR is collecting and curating biomedical human data from a new sub-orbital research flight and is open to more space life science/biomedical submissions from the international and commercial sectors. OSDR also recently began a collaboration with the European Space Agency (ESA) to collect and curate >200 terabytes of human and model organism data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics and ~50 physiological-phenotypic-imaging assay data types. Tools available for OSDR users include: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, and 3) a Multi-study visualization tool which enables users to look across and combine ‘omics datasets. There are ~600 volunteer OSDR Analysis Working Group (AWG) members providing feedback on scientific data/metadata standards and collaborating to mine-reuse OSDR in research. OSDR/GeneLab has enabled ~60 publications reusing data as of October 2023.

space biology↗

Pypromice: A Python Package for Processing Automated Weather Station Data

The pypromice Python package is for processing and handling observation datasets from automated weather stations (AWS). It is primarily aimed at users of AWS data from the Geological Survey of Denmark and Greenland (GEUS), which collects and distributes in situ weather station observations to the cryospheric science research community. Functionality in pypromice is primarily handled using two key open-source Python packages, xarray (Hoyer & Hamman, 2017) and pandas (The pandas development team, 2020). A defined processing workflow is included in pypromice for transforming original AWS observations (Level 0, L0) to a usable, CF-convention-compliant dataset (Level 3, L3) (Figure 1). Intermediary processing levels (L1,L2) refer to key stages in the workflow, namely the conversion of variables to physical measurements and variable filtering (L1), cross-variable corrections and user-defined data flagging and fixing (L2), and derived variables (L3). Information regarding the station configuration is needed to perform the processing, such as instrument calibration coefficients and station type (one-boom tripod or two-boom mast station design, for example), which are held in a toml configuration file. Two example configuration files are provided with pypromice , which are also used in the package’s unit tests. More detailed documentation of the AWS design, instrumentation, and processing steps are described in Fausto et al. (2021).

pypromice↗

Fostering Open Science Inclusiveness for Interdisciplinary Users of Earth Observations

The term Open Science is subject to a variety of interpretations because of a key (and useful) ambiguity in the meaning of “Open”. Open in the sense of Transparency enables more trust in science research by making the details of the scientific process visible and accessible to anyone. “Open” in the sense of Inclusiveness enables more scientists from other disciplines to participate in research in a given discipline, thus producing more interdisciplinary research. Data Systems can play a major role in enabling Open (Inclusive) Science by making it easier for users from other disciplines to work with data within a given discipline. This is challenging for Earth Observation datasets, most of which are the product of advanced instrumentation and sophisticated, specialized variable retrieval algorithms and code. Serving the “extra-disciplinary”communities begins with simple things, like accessible, readable data documentation with adequate scaffolding. But just as important is provisioning Analysis-Ready data that does not require expert pre-processing. Disciplines also often have dominant toolsets, such as R in the biomass community or GIS in many applications communities. Ensuring that EO data are easy to use in the tools favored in other communities will enable more interdisciplinary research. Ideally, interdisciplinary research also benefits from scientists with different domain expertise. Platforms and frameworks that facilitate frictionless collaboration with discipline experts, together with capacity building efforts in those external disciplines also improve the inclusiveness aspect of Open Science. In short, Open Science is at root a way of thinking about how users from diverse discipline can best access and use data and services from a particular discipline.

Christopher Lynnes↗

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert↗

Dark Energy Survey: Implications for cosmological expansion models from the final DES baryon acoustic oscillation and supernova data

The Dark Energy Survey (DES) recently released the final results of its two principal probes of the expansion history: Type Ia supernovae (SNe) and baryonic acoustic oscillations (BAO). In this paper, we explore the cosmological implications of these data in combination with external cosmic microwave background (CMB), big bang nucleosynthesis (BBN), and age-of-the-Universe information. The BAO measurement, which is ∼ 2 σ away from Planck ’s Λ CDM predictions, pushes for low values of Ω m compared to Planck, in contrast to SN which prefers a higher value than Planck. We identify several tensions among datasets in the Λ CDM model that cannot be resolved by including either curvature ( k Λ CDM ) or a constant dark energy equation of state ( w CDM ). By combining BAO + SN + CMB despite these mild tensions, we obtain Ω k = - 5.5 - 4.2 + 4.6 × 10 - 3 in k Λ CDM , and w = - 0.94 8 - 0.027 + 0.028 in w CDM . In w CDM , BAO and SN push again in different directions of parameter space, favoring, respectively, w < - 1 and w > - 1 . If we open the parameter space to w 0 w a CDM [where the equation of state of dark energy varies as w ( a ) = w 0 + ( 1 - a ) w a ], all the datasets are mutually more compatible, and we find concordance in the [ w 0 > - 1 , w a < 0 ] quadrant, with BAO pushing for w a < 0 and SN for [ w 0 > - 1 , w a < 0 ] . For DES BAO and SN in combination with Planck -CMB, we find a 3.2 σ deviation from Λ CDM , with w 0 = - 0.67 3 - 0.097 + 0.098 , w a = - 1.3 7 - 0.50 + 0.51 , a Hubble constant of H 0 = 67.8 1 - 0.86 + 0.96 km s - 1 Mpc - 1 , and an abundance of matter of Ω m = 0.310 9 - 0.0099 + 0.0086 . For the combination of all the background cosmological probes considered (including CMB’s angular acoustic scale θ ⋆ ), we still find a deviation of 2.8 σ from Λ CDM in the w 0 - w a plane. Assuming a minimal neutrino mass, this work provides tentative evidence for non- Λ CDM physics, which is consistent with recent claims in support of evolving dark energy, or a source of unknown systematics.

79 ASTRONOMY AND ASTROPHYSICS↗

A new database website for nuclear level densities

We introduce a new open-access, web-based database (http://nld.ascsn.net), Current Archive of Nuclear Density of Levels (CANDL), that hosts experimental nuclear level density (NLD) datasets from a variety of techniques and energy ranges. Built using the Dash framework in Python, the database is designed to be interactive and user-friendly, allowing researchers to search, visualize, fit, and export NLD data with minimal effort. This resource includes data extracted from evaporation spectra, Oslo method variants, and other experimental techniques that cover excitation energies beyond the neutron resonance region. The database supports on-the-fly fitting with two widely-used phenomenological models—the Constant Temperature (CT) model and the Back-Shifted Fermi Gas (BSFG) model—selected for their simplicity and computational efficiency. Future versions aim to include additional datasets and model types, as well as easy-to-use interfaces to data science techniques. Here, this platform offers a vital tool for the nuclear physics, astrophysics, medicine, and reactor design communities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Atmospheric Correction Inter-Comparison eXercise

The Atmospheric Correction Inter-comparison eXercise (ACIX) is an international initiative with the aim to analyse the Surface Reflectance (SR) products of various state-of-the-art atmospheric correction (AC) processors. The Aerosol Optical Thickness (AOT) and Water Vapour (WV) are also examined in ACIX as additional outputs of AC processing. In this paper, the general ACIX framework is discussed; special mention is made of the motivation to initiate the experiment, the inter-comparison protocol, and the principal results. ACIX is free and open and every developer was welcome to participate. Eventually, 12 participants applied their approaches to various Landsat-8 and Sentinel-2 image datasets acquired over sites around the world. The current results diverge depending on the sensors, products, and sites, indicating their strengths and weaknesses. Indeed, this first implementation of processor inter-comparison was proven to be a good lesson for the developers to learn the advantages and limitations of their approaches. Various algorithm improvements are expected, if not already implemented, and the enhanced performances are yet to be assessed in future ACIX experiments.

Processors inter-comparison↗

Spaceflight Biospecimen Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have flown non-human biological experiments in space to understand the effects of spaceflight and address potential biological hazards. Sending organisms into space is a costly endeavor which makes space-flown biological specimens a valuable resource. To enable maximum scientific return, samples not required by the Principal Investigators are harvested and collected mostly by NASA’s Space Biology Biospecimen Sharing Program. These specimens are collected according to well-established SOPs that maintain quality and integrity. The specimens are then preserved, archived, and made available to the international scientific community through NASA’s Institutional Scientific Collection (ISC) at Ames Research Center (ARC). The ISC-ARC biospecimens and descriptive metadata are findable and accessible for request through the Life Sciences Data Archive (LSDA). The NASA ISC-ARC currently stores over 32,000 specimens from Shuttle, International Space Station, and ground-based investigations (spaceflight analog experiments involving either hindlimb unloading, centrifugation, or partial weight-bearing study designs). Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. The specimens include tissues from many physiological systems including musculoskeletal, neurosensory, reproductive, respiratory, circulatory, and digestive. Tissues are stored at -80°C, -20°C, +4°C, or ambient and preserved in various fixatives. Descriptive metadata is available for all samples. Historically, these tissues have been used for a wide range of analyses, including histology, genomics, and transcriptomics. Plans are underway to expand the ISC-ARC beyond the mostly-rodent contents, to include a space-relevant microbial culture collection including bacteria, fungi, and yeast. This expansion of the ISC-ARC will now involve identifying and standardizing best practices for microbial curations. To ensure safe long-term storage of microbial isolates, a microbiology laboratory will be dedicated for identification, cell culture, and lyophilization. Awarding of tissue to public science investigators has resulted in 33 publications since 2011, with 48 requests being submitted since 2016. Of note, NASA GeneLab has been awarded ISC-ARC biospecimens in the past few years. GeneLab processes the biospecimens to generate various levels of ‘omics’ data, which are published on GeneLab’s open access online platform for bioinformatics analysis and visualization. This has helped a systems biology community grow around the processed-biospecimens’ datasets, resulting in many new publications and insights. Websites: https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://lsda.jsc.nasa.gov/Biospecimen

Ryan T. Scott↗

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Reversed-direction 2-point modelling applied to divertor conditions in DIII-D *

Abstract A predictive form of the extended 2-point model known as the ‘reverse 2-point model’, Rev2PM, is applied to a range of detachment levels in the open lower divertor of DIII-D, showing that the experimentally measured electron temperature ( T e ) and pressure ( p e ) at the divertor entrance can be calculated within 50% from target measurements, if and only if a posteriori corrections for convective heat flux are included in the model. Unlike the standard 2-point model, the Rev2PM calculates upstream scrape-off layer (SOL) quantities (such as separatrix T e and p e ) from target conditions (such as T e and parallel heat flux), with volumetric power and momentum losses depending solely on target T e . The Rev2PM is tested against a database of DIII-D inter-ELM divertor Thomson scattering measurements, built from a series of 6 MW, 1.3 MA, LSN H-mode discharges with varied main ion density, drift direction, and nitrogen puffing rate. Measured target T e ranged from 0.4–25 eV over this database, and upstream T e ranged from 5–60 eV. Poor agreement is found between upstream measurements and Rev2PM calculations that assume purely conductive parallel heat transport. However, introducing a posteriori corrections to account for convective heat transport brings the Rev2PM calculations within 50% of the measured upstream values across the dataset. These corrections imply that up to 99% of the parallel heat flux is carried by convection in detached conditions in the DIII-D open lower divertor, though further work is required to assess any potential dependencies on device size or divertor closure.

2-point model↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗