Search NASA⌕ Search

SEARCH · Search NASA

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Development of WMS Capabilities to Support NASA Disasters Applications and App Development

During the last year several significant disasters have occurred such as Superstorm Sandy on the East coast of the United States, and Typhoon Bopha in the Phillipines, along with several others. In support of these disasters NASA's Short‐term Prediction Research and Transition (SPoRT) Center delivered various products derived from satellite imagery to help in the assessment of damage and recovery of the affected areas. To better support the decision makers responding to the disasters SPoRT quickly developed several solutions to provide the data using open Geographical Information Service (GIS) formats. Providing the data in open GIS standard formats allowed the end user to easily integrate the data into existing Decision Support Systems (DSS). Both Tile Mapping Service (TMS) and Web Mapping Service (WMS) were leveraged to quickly provide the data to the end‐user. Development of the deliver methodology allowed quick response to rapidly developing disasters and enabled NASA SPoRT to bring science data to decision makers in a successful research to operations transition.

Bell, Jordan R.↗

Development of WMS Capabilities to Support NASA Disasters Applications and App Development

During the last year several significant disasters have occurred such as Superstorm Sandy on the East coast of the United States, and Typhoon Bopha in the Phillipines, along with several others. In support of these disasters NASA's Short‐term Prediction Research and Transition (SPoRT) Center delivered various products derived from satellite imagery to help in the assessment of damage and recovery of the affected areas. To better support the decision makers responding to the disasters SPoRT quickly developed several solutions to provide the data using open Geographical Information Service (GIS) formats. Providing the data in open GIS standard formats allowed the end user to easily integrate the data into existing Decision Support Systems (DSS). Both Tile Mapping Service (TMS) and Web Mapping Service (WMS) were leveraged to quickly provide the data to the end‐user. Development of the deliver methodology allowed quick response to rapidly developing disasters and enabled NASA SPoRT to bring science data to decision makers in a successful research to operations transition.

Bell, Jordan R.↗

Development of Web Mapping Service Capabilities to Support NASA Disasters Applications / App Development

During the last year several significant disasters have occurred such as Superstorm Sandy on the East coast of the United States, and Typhoon Bopha in the Phillipines, along with several others. In support of these disasters NASA's Short‐term Prediction Research and Transition (SPoRT) Center delivered various products derived from satellite imagery to help in the assessment of damage and recovery of the affected areas. To better support the decision makers responding to the disasters SPoRT quickly developed several solutions to provide the data using open Geographical Information Service (GIS) formats. Providing the data in open GIS standard formats allowed the end user to easily integrate the data into existing Decision Support Systems (DSS). Both Tile Mapping Service (TMS) and Web Mapping Service (WMS) were leveraged to quickly provide the data to the end‐user. Development of the deliver methodology allowed quick response to rapidly developing disasters and enabled NASA SPoRT to bring science data to decision makers in a successful research to operations transition.

Burks, Jason E.↗

Development of Web Mapping Service Capabilities to Support NASA Disasters Applications/App Development

During the last year several significant disasters have occurred such as Superstorm Sandy on the East coast of the United States, and Typhoon Bopha in the Phillipines, along with several others. In support of these disasters NASA's Short‐term Prediction Research and Transition (SPoRT) Center delivered various products derived from satellite imagery to help in the assessment of damage and recovery of the affected areas. To better support the decision makers responding to the disasters SPoRT quickly developed several solutions to provide the data using open Geographical Information Service (GIS) formats. Providing the data in open GIS standard formats allowed the end user to easily integrate the data into existing Decision Support Systems (DSS). Both Tile Mapping Service (TMS) and Web Mapping Service (WMS) were leveraged to quickly provide the data to the end‐user. Development of the deliver methodology allowed quick response to rapidly developing disasters and enabled NASA SPoRT to bring science data to decision makers in a successful research to operations transition.

Burks, Jason E.↗

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications↗

NASA's Earth Science Data Systems Standards Process Experiences

NASA has impaneled several internal working groups to provide recommendations to NASA management on ways to evolve and improve Earth Science Data Systems. One of these working groups is the Standards Process Group (SPC). The SPG is drawn from NASA-funded Earth Science Data Systems stakeholders, and it directs a process of community review and evaluation of proposed NASA standards. The working group's goal is to promote interoperability and interuse of NASA Earth Science data through broader use of standards that have proven implementation and operational benefit to NASA Earth science by facilitating the NASA management endorsement of proposed standards. The SPC now has two years of experience with this approach to identification of standards. We will discuss real examples of the different types of candidate standards that have been proposed to NASA's Standards Process Group such as OPeNDAP's Data Access Protocol, the Hierarchical Data Format, and Open Geospatial Consortium's Web Map Server. Each of the three types of proposals requires a different sort of criteria for understanding the broad concepts of "proven implementation" and "operational benefit" in the context of NASA Earth Science data systems. We will discuss how our Standards Process has evolved with our experiences with the three candidate standards.

Ullman, Richard E.↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

Using Big Data Technologies with Earth Science Data in HDF5: HDF5 Scalable Solutions

HDF5 (Hierarchical Data Format 5) is open-source, high-performance software that consists of an abstract data model, library, and fileformat used for storing and managing extremely large and/or complex data collections. NASA Earth Observing System (EOS) Data and Information Systems use HDF5 as an archival format to store remote sensing data from EOS satellites. HDF5 is also used to store other types of Geoscience and Strophysical data, e.g., seismic data and data from Low-Frequency Array (LOFAR) radio telescopes. Data stored in HDF5 has reached tens of petabytes and is growing at an accelerated rate.With the growing amout of HDF5 Earth Science data to analyze and process, scientists need to adopt big data technologies including new storage paradigms such as cloud and object storage. To run models and perform data analysis they also need to utilizied efficient and diverse ways to access data, from high-performance computing's (HPC) Message Passing Interface (MPI) I/O and deep memory hierarchies (DMH) to non-HPC frameworks such as Apache Hadoop, Spark, and Drill. The HDF Group continually works to enable usage of big data technologies in HDF software.

Knox, Larry↗

Open Source GIS Connectors to NASA GES DISC Satellite Data

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) houses a suite of high spatiotemporal resolution GIS data including satellite-derived and modeled precipitation, air quality, and land surface parameter data. The data are valuable to various GIS research and applications at regional, continental, and global scales. On the other hand, many GIS users, especially those from the ArcGIS community, have difficulties in obtaining, importing, and using our data due to factors such as the variety of data products, the complexity of satellite remote sensing data, and the data encoding formats. We introduce a simple open source ArcGIS data connector that significantly simplifies the access and use of GES DISC data in ArcGIS.

user↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching. The use of health countermeasures and biomonitoring systems for space missions are required to counteract space health hazards and to support life to thrive in deep space (e.g., humans, animals, plants, crops; entire ecosystems within spacecrafts/habitats/spacesuits). The development of these mission components will be highly dependent on our understanding of basic biological and health responses to myriad space hazards (ionizing radiation, altered gravitational fields, altered day-night cycles, confined isolation, hostile-closed environments, distance-duration from Earth, planetary dust-regolith, and extreme temperatures/atmospheres). The fast-growing array of space biological and mission telemetry data, which in the past was simply archived after minimal analysis, holds great potential once applied to these mission challenges if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its multi-hierarchical, multi-modal, and heterogenous nature (molecular, cellular, tissue, organ, whole organism, behavior, ecosystem, microbiome; tabular, omics, imaging, video, biospecimen, environmental physical-chemical telemetry). This session focuses on current approaches in this domain such as: making space biological data FAIR (findable, accessible, interoperable, reusable), effective data ingestion/dissemination, observational versus experimental data, Open Science collaborations, data analysis techniques, AI/ML/knowledge graph/modeling methods, and data integration/discovery tools.

open science↗

HEASARC - The High Energy Astrophysics Science Archive Research Center

The High Energy Astrophysics Science Archive Research Center (HEASARC) is NASA's archive for high-energy astrophysics and cosmic microwave background (CMB) data, supporting the broad science goals of NASA's Physics of the Cosmos theme. It provides vital scientific infrastructure to the community by standardizing science data formats and analysis programs, providing open access to NASA resources, and implementing powerful archive interfaces. Over the next five years the HEASARC will ingest observations from up to 12 operating missions, while serving data from these and over 30 archival missions to the community. The HEASARC archive presently contains over 37 TB of data, and will contain over 60 TB by the end of 2014. The HEASARC continues to secure major cost savings for NASA missions, providing a reusable mission-independent framework for reducing, analyzing, and archiving data. This approach was recognized in the NRC Portals to the Universe report (2007) as one of the HEASARC's great strengths. This poster describes the past and current activities of the HEASARC and our anticipated developments in coming years. These include preparations to support upcoming high energy missions (NuSTAR, Astro-H, GEMS) and ground-based and sub-orbital CMB experiments, as well as continued support of missions currently operating (Chandra, Fermi, RXTE, Suzaku, Swift, XMM-Newton and INTEGRAL). In 2012 the HEASARC (which now includes LAMBDA) will support the final nine-year WMAP data release. The HEASARC is also upgrading its archive querying and retrieval software with the new Xamin system in early release - and building on opportunities afforded by the growth of the Virtual Observatory and recent developments in virtual environments and cloud computing.

Smale, Alan P.↗

Connecting Users and Applications with Po.daac Hosted GHRSST Data

The 80+ GHRSST public datasets represent a rich resource for sea surface temperature research and applications given their time series length, resolution, spatial coverage, varying measurement types and processing levels, and availability in the full spectrum of PO.DAAC tools and services ecosystem. The PO.DAAC has created a publicly accessible recipe suite for the user community to perform straightforward yet powerful computations on GHRSST data using python recipes, Jupyter notebooks, R, Matlab, and the NCO programming language. These recipes include numerical computations for regional and global SST trends, anomaly derivations, EOF analysis, climate signal reproduction, and ocean phenology. For example, one recipe reproduces a famous SST based warming figure from the Fourth National Climate Assessment (USA) while another focuses on quantifying the regional changes in ocean SST phenology. Most are python-based while some contain hybrid calls and leverage the NCO programming interface too. All are available on the PO.DAAC user forum (https://podaac.jpl.nasa.gov/forum/) and/or via the open source NASA GitHub repository (https://github.com/nasa/podaac_tools_and_services). Several are available in the Jupyter notebook framework including podaacypy (https://github.com/nasa/podaacpy), a recipe for GHRSST granule metadata discovery and application, and more recently a Jupyter notebook developed to support data analysis and visualization of a cloud-based Zarr formatted Level 4 MUR dataset in the AWS Open Data Registry. Throughout the summer of 2020, the PO.DAAC intends to add and migrate more of its numerical recipes to the Jupyter notebook framework and publish them on its open source GitHub repository.

Gentemann, Chelle↗

Formal Provenance Representation of the Data and Information Supporting the National Climate Assessment

The Global Change Information System (GCIS) provides a framework for the formal representation of structured metadata about data and information about global change. The pilot deployment of the system supports the National Climate Assessment (NCA), a major report of the U.S. Global Change Research Program (USGCRP). A consumer of that report can use the system to browse and explore that supporting information. Additionally, capturing that information into a structured data model and presenting it in standard formats through well defined open inter- faces, including query interfaces suitable for data mining and linking with other databases, the information becomes valuable for other analytic uses as well.

Provenance↗

The Deep-Time Digital Earth program: data-driven discovery in geosciences

Current barriers hindering data-driven discoveries in deep-time Earth (DE) include: substantial volumes of DE data are not digitized; many DE databases do not adhere to FAIR (findable, accessible, interoperable and reusable) principles; we lack a systematic knowledge graph for DE; existing DE databases are geographically heterogeneous; a significant fraction of DE data is not in open-access formats; tailored tools are needed. These challenges motivate the Deep-Time Digital Earth (DDE) program initiated by the International Union of Geological Sciences and developed in cooperation with national geological surveys, professional associations, academic institutions and scientists around the world. DDE’s mission is to build on previous research to develop a systematic DE knowledge graph, a FAIR data infrastructure that links existing databases and makes dark data visible, and tailored tools for DE data, which are universally accessible. DDE aims to harmonize DE data, share global geoscience knowledge and facilitate data-driven discovery in the understanding of Earth’s evolution.

Chengshan Wang↗

Biomass Harmonization and SAR Analysis with the Multi-mission Algorithm and Analysis Platform (MAAP)

The Multi‐mission Algorithm and Analysis Platform (MAAP) is a collaborative effort between NASA and the European Space Agency (ESA) to support above ground biomass (AGB) research in an open science framework. MAAP brings together relevant data, algorithms, and computing capabilities in a common cloud environment to address the challenges of sharing and processing data from field, airborne and satellite measurements. MAAP was publicly released in October 2021, providing computing capabilities co-located with the data, a collaborative coding and analysis environment, and a set of interoperable tools and algorithms developed to support the estimation and visualization of data. MAAP has allowed scientists from both North America and Europe to collaborate on the generation and analysis/visualization of data derived from multiple, discipline-adjacent missions in an open, collaborative environment that has reached beyond traditional scientific investigation. MAAP has been used to support multiple scientific activities. To date, existing LiDAR data from multiple platforms has been calibrated with field measurements and combined for more comprehensive and accurate estimates of above ground biomass AGB; these LiDAR platforms include airborne (e.g. LVIS), the International Space Station (NASA’s Global Ecosystem Dynamics Investigation (GEDI), and satellites (e.g. ICESat-2). The current challenge is to effectively and seamlessly combine the aforementioned LiDAR-based data with new data sources such as P-band RADAR from ESA’s upcoming BIOMASS mission, existing ESA Sentinel-1 C-band SAR, and the 30 PB/yr of high cadence global coverage L-band SAR data from the upcoming NASA-ISRO SAR (NISAR) mission. Recent analysis using MAAP merged ICESat-2 and optical data (Harmonized Landsat Sentinel) produced the most comprehensively precise estimate of boreal-wide AGB to date. Another effort using MAAP is the production and open distribution of global comparisons of AGB map estimates, including from ICESat-2 and GEDI, to bolster stakeholder uptake for policy applications. These map estimates will feed into the Intergovernmental Panel on Climate Change (IPCC) database, likely aiding the next Global Carbon Stocktake of the UNFCCC. Furthermore, the biomass retrieval intercomparison exercise BRIX-2 could benefit from the MAAP providing standardized test cases (based on airborne campaign and spaceborne data) allowing the community to develop and apply retrieval algorithms based on these test cases, while forthcoming SAR data training curricula could also use the MAAP as a teaching and learning platform. The MAAP is meeting the challenges inherent in international, open science collaboration and large scale computing with a platform that is entirely open source and cloud native, using open standards for data access, manipulation, protocols, and formats. The MAAP data system consists of a dedicated data store whose data is indexed in an online catalog conforming to established metadata, application programmatic interfaces (APIs), and service interface standards, using an implementation of the open sourced NASA Common Metadata Repository. Federation of user identities allows users from either NASA or ESA to access and consume services from the other using a unified metadata catalog for the data utilized across the ESA and NASA MAAP platforms. Similarly, we are exploring how to increase interoperability to achieve a common approach to packaging, orchestrating and executing algorithms, with interoperable access to data for subsetting, fast browse, and cloud-optimized access, all using interoperable standards such as those from the Open Geospatial Consortium (OGC). Designed for interoperability, ESA and NASA utilize a common architecture for the software platform. It provides a cloud-based algorithm development environment (ADE) that enables scientists to develop algorithms collaboratively with access to the MAAP data catalog as well as other data archives. MAAP provides an Eclipse Che-based ADE supporting both Python and R languages, popular in this biomass community. Algorithms developed and containerized within the ADE can be deployed to run to thousands of computational nodes in the MAAP’s data processing system (DPS), dramatically speeding up processing and giving scientists a rapid, iterative turnaround of results. NASA’s implementation of the DPS is based on the Hybrid Science Data System (HySDS) framework, used by NASA flight projects to produce Earth science standard products.

cloud computing↗

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗