Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CSW Best Practices

During the development of the CMR (Common Metadata Repository) (CMR) for the Earth Observing System Data and Information System (EOSDIS), CSW (Catalog Service for the Web) a number of best practices came to light. Given that the ESIP (Earth Science Information Partners) Discovery Cluster is committed to interoperability and standards in earth data discovery this seemed like a convenient moment to provide Best Practices to the organization in the same way we did for OpenSearch for this widely-used standard.

CMR↗

NASA Earth Observing System Data and Information System (EOSDIS): A U.S. Network of Data Centers Serving Earth Science Data: A Network Member of ICSU WDS

NASA's Earth Observing System Data and Information System (EOSDIS) has been in operation since August 1994, and serving a diverse user community around the world with Earth science data from satellites, aircraft, field campaigns and research investigations. The ESDIS Project, responsible for EOSDIS is a Network Member of the International Council for Sciences (ICSU) World Data System (WDS). Nine of the 12 Distributed Active Archive Centers (DAACs), which are part of EOSDIS, are Regular Members of the ICSUWDS. This poster presents the EOSDIS mission objectives, key characteristics of the DAACs that make them world class Earth science data centers, successes, challenges and best practices of EOSDIS focusing on the years 2014-2016, and illustrates some highlights of accomplishments of EOSDIS. The highlights include: high customer satisfaction, growing archive and distribution volumes, exponential growth in number of products distributed to users around the world, unified metadata model and common metadata repository, flexibility provided to uses by supporting data transformations to suit their applications, near-real-time capabilities to support various operational and research applications, and full resolution image browse capabilities to help users select data of interest. The poster also illustrates how the ESDIS Project is actively involved in several US and international data system organizations.

Earth Science↗

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, to foster collaborations with equity. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within the Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate with community members and set new standards for space-relevant data and metadata. The AWGs welcome any ASGSR members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are encouraged to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR↗

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, and fostering collaborations with equity. 2023 is the ‘Year of Open Science’, and NASA has a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within NASA’s Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate and set new standards for space-relevant data and metadata. The AWGs welcome any ASPB members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are invited to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR↗

Storage of Physical Sample Metadata in the Astrobiology Habitable Environments Database (AHED)

The National Aeronautics and Space Administration has begun an effort to store, curate, and publish information about physical samples collected and analyzed in conjunction with NASA-funded astrobiology research. Astrobiology is a multidisciplinary area of scientific research being conducted by collaborating teams of biologists, chemists, geologists, atmospheric scientists, oceanographers, astrophysicists, astronomers, and other specialists. Astrobiology studies the origin, evolution, and distribution of life in the Universe. NASA uses the results of astrobiology research to focus its future missions on targets of opportunity for the discovery of life off Earth. Astrobiology researchers conduct both field-based and laboratory-based research, during which physical samples are collected, processed, and catalogued. The cataloguing practices employed by different teams of astrobiologists vary widely, and there are no specific standards available to guide the collection and recording of astrobiology sample data. The disparity in data collection approaches and the lack of a centralized sample repository makes it difficult for astrobiology teams to share data and benefit from resultant synergies.To facilitate data sharing within the astrobiology community, NASA is developing a prototype database the Astrobiology Habitable Environments Database (AHED) and an associated set of data collection templates. The database will store information about samples, along with associated measurements and analyses, including information about biological cultures enriched or isolated from samples, and the results of analyses performed on the samples (e.g., via spectrography, microscopy, etc.). In addition, the system will store contextual information about field sites where samples were collected, the instruments or equipment used for analysis, and people and institutions involved in their collection. AHED is being implemented on top of Open Data Repository's Data Publisher [1], an open source software platform for the publication of scientific datasets. The data collection templates under development represent an initial attempt to propose a set of metadata for capture and storage within AHED. The design of these templates is being conducted by a consolidated group of astrobiologists from active research teams at NASA Ames Research Center, assisted by data science and software engineering specialists. These initial templates must be vetted with the broader astrobiology community through a defined process to ensure that they meet community needs. Each template captures a different type of data collection record. For each template, we are developing a list of fields to be captured, including a set of required entry fields, a set of recommended but optional fields, and a set of discretionary fields. A datatype selected from a variety of text and numeric types is specified for each field. Included is a 'choice' type that restricts user input to an enumerated list of values. Many of the fields and field values capture information of particular interest to the astrobiology community, and are intended to facilitate search and retrieval of relevant data across multiple datasets.

Keller, Rich↗

NASA's Next Generation of Atmospheric Data Science

The Multi-Angle Imager for Aerosols (MAIA) and the Tropospheric Emission: Monitoring of Pollution(TEMPO) are NASA’s next-generation satellite missions for air quality monitoring. These missions will produce high-quality, high-resolution air quality data to support cross-displinary research. The MAIA mission is collaborating with health science researchers and epidemiologists to study the impacts of air quality on health outcomes. TEMPO aims to improve our understanding of tropospheric air pollution chemistry and our ability to make predictions about air quality and climate forcing. TEMPO will offer hourly measurements of tropospheric ozone, aerosols, and clouds focused on North America at high-spatial resolution, while MAIA will produce high-resolution measurements of speciated particulate matter targeting densely populated cities around the globe. Data from these missions will help improve our understanding of the sources, types, and interactions among the aerosols and trace gases that are polluting Earth’s atmosphere, as well as our understanding of the impact of air pollution on pollution on a wide range of important areas including human health, agriculture, weather, and climate change. The challenges of cross-disciplinary research, computationally expensive multi-variate analyses, and high-resolution data at both local and global scales are driving substantial changes across all of NASA’s Distributed Active Archive Centers (DAACs). High resolution data at scales such these requires a new approach to data ingest, archive, and publication. Like other NASA DAACs, the Atmospheric Science Data Center (ASDC), the DAAC that will be responsible for publishing MAIA and TEMPO data products has historically archived and distributed data on premise. DAACs of the future will archive and distribute data in the cloud, enabling them to remake themselves as research-focused data centers that will support on-demand, data-intensive computations for highly accurate retrospective analyses and predictions. Under the new paradigm, data formats and metadata must support on-demand spatial and temporal sub-setting, as well as other data transformation services such as re-gridding and re-sampling. This presentation will discuss work being done to address data formatting and metadata requirements in this dynamic new environment. In addition to the changes in data stewardship practices at the ASDC, the increased focus on supporting scientific research is driving changes in the relationship between DAACs and researchers. While the ASDC will continue to provide first rate data management and stewardship, it is increasingly focused on serving as a partner not only to the science teams that gather and produce the data it publishes, but to the researchers that use that data.

Beth Huffer↗

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena↗

Interoperability and Other Aspects of Guiding Data Producers for the Benefit of End Users

The purpose of this paper is to discuss how the Climate and Forecast (CF) Metadata Conventions and netCDF standard have influenced the recommendations and guidance provided to producers of data products based on NASA’s Earth observations. It has been long-recognized that interoperable datasets and use of standards and conventions are beneficial to the users of these datasets, especially those who make use of multiple datasets for their research and applications. The Dataset Interoperability Working Group (DIWG), one of NASA’s Earth Science Data System Working Groups (ESDSWGs), was established in 2013, and has developed and published many recommendations. The Data Product Development Guide (DPDG) Working Group, established in 2018 as another of the ESDSWGs, has published a DPDG for Data Producers and a Quick Start Guide, incorporating guidance from many sources, including the recommendations from the DIWG. The DPDG includes recommendations regarding data formats (prominently netCDF-4) and metadata based primarily on the CF Metadata Conventions and the Attribute Convention for Data Discovery (ACDD). In early 2023, it was decided that the Resource Center for Data Producers (RCDP) Working Group be established as another ESDSWG, with the goals of providing all the information relevant and helpful for data producers via an easily accessible website, and of recommending how the DPDG and QSG could be maintained as living documents, given the rapidly changing technologies, and the need for incorporating the experience and feedback from the users of these documents.

Data product development↗

Perspectives on Data Reproducibility and Replicability in Paleoclimate and Climate Science

This paper summarizes the current state of reproducibility and replicability in the fields of climate and paleoclimate science, including brief histories of their development and applications in climate science, new and recent approaches towards improvement of reproducibility and replicability, and challenges. Recommendations for addressing those challenges include: development of searchable, auto-updated, interlinked, multi-archive public paleoclimate repositories for raw and processed digital datasets; cross-center standardized code base cases, improved data storage techniques, and a focus on replicability for climate simulation storage and access; and support of the development and community awareness of findable, accessible, interoperable and reusable (FAIR) principles by funding agencies and publishers. This paper is largely based on the May 2018 presentations of a panel of researchers to the Committee on Reproducibility and Replicability in Science, part of the National Academies of Science, Engineering, and Medicine. The commentary and recommendations made here are in alignment with those of its Consensus Study Report on Reproducibility and Replicability in Science (2019).

data repositories↗

Simplifying Analysis of Hierarchical HDF5 and NetCDF4 Files with Xarray-Datatree

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. EOSDIS data are stored in formats that are well supported by the Earth Science community. These formats include the Hierarchical Data Format (HDF), with derivative flavors such as HDF-5 and the Network Common Data Format (NetCDF-4). The HDF specification allows for a directory-like hierarchy within a single file, known as "groups". Observational data and associated metadata within a single file can be distributed amongst multiple internal groups, which can also be nested to multiple levels. Working with datasets that have a group hierarchical structure can be difficult because of the nested structure of groups. Widely used packages, such as xarray, have data models that do not accommodate the hierarchical structure within HDF files, requiring users to traverse the file and open different HDF groups as separate, unrelated objects. Xarray-datatree is a Python package developed to solve the difficulty of traversing HDFs with a hierarchical group structure by creating a tree-like hierarchical data structure in xarray. The tree-like structure allows each group to be accessed once a DataTree object is instantiated. The migration of xarray-datatree into the xarray core library will reduce barriers to accessing Earth science data by eliminating the need to understand and traverse the specific hierarchy of a grouped HDF file.

Eni Awowale↗

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions rely on plants and crops for crew and ecosystem health. Access to space plant data enables scientists to gain a deeper understanding of biological responses to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, and altered photoperiods. Open Science is the practice of making research available to all, while respecting diverse cultures, fostering collaborations with equity. 2023 is the ‘Year of Open Science’, and NASA has a 5-year Transform to Open Science (TOPS) mission designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) developed by NASA’s Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. OSDR started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository (GeneLab), providing detailed metadata on investigation, sample, and assay levels. Today, GeneLab hosts 62 plant datasets which have led to 5 published peer-reviewed meta-analysis publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate and set new standards for space-relevant data and metadata. The AWGs are welcoming any ASPB members interested in providing plant expertise for space biology. The addition of ALSDA to OSDR is also expanding analysis capability beyond ‘omics. Now is the time to get involved as a Subject Matter Expert as we establish the framework for modern plant data archiving through the AWGs. Investigators are invited to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR↗

Data Sharing in Radiobiology; Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally „Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Data Sharing in Radiation Biology: Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally "Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Simple, Script-Based Science Processing Archive

The Simple, Scalable, Script-based Science Processing (S4P) Archive (S4PA) is a disk-based archival system for remote sensing data. It is based on the data-driven framework of S4P and is used for data transfer, data preprocessing, metadata generation, data archive, and data distribution. New data are automatically detected by the system. S4P provides services such as data access control, data subscription, metadata publication, data replication, and data recovery. It comprises scripts that control the data flow. The system detects the availability of data on an FTP (file transfer protocol) server, initiates data transfer, preprocesses data if necessary, and archives it on readily available disk drives with FTP and HTTP (Hypertext Transfer Protocol) access, allowing instantaneous data access. There are options for plug-ins for data preprocessing before storage. Publication of metadata to external applications such as the Earth Observing System Clearinghouse (ECHO) is also supported. S4PA includes a graphical user interface for monitoring the system operation and a tool for deploying the system. To ensure reliability, S4P continuously checks stored data for integrity, Further reliability is provided by tape backups of disks made once a disk partition is full and closed. The system is designed for low maintenance, requiring minimal operator oversight.

Lynnes, Christopher↗

A Standard Reference Model for Data Archives

An implementable Data Archive Architecture is being developed for trusted digital repositories based on the Reference Model for an Open Archival Information System (OAIS) – ISO 14721. A set of interoperable protocols and interface specifications are planned that will offer capabilities for accessing, merging, and re-using data, both within and across the operational boundaries of trustworthy digital repositories. The model will also provide support for the fundamental scientific need to verify the reproducibility of results. This standards development task is being performed by the Data Archive Interoperability (DAI) working group within the Consultative Committee for Space Data Systems (CCSDS). The architecture integrates concepts from the OAIS Reference Model, the ISO/IEC 11179 Metadata Registry (MDR) standard, the CCSDS Reference Architecture for Space Information Management (RASIM), the proposed draft recommended practice document, Information Preparation to Enable Long Term Use (IPELTU), and three decades of digital repository development for science research.

Ambacher, Bruce↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler↗

Integrating Multi-agency Data Products in a Cloud-based Platform for Streamlined Discovery, Visualization, and Use

Earth science data users almost always have an interest in utilizing geospatial data from multiple agencies. As computing capability and cloud-based infrastructures accelerate the pace at which scientific research can be done, there is a growing need to enable search, discovery, and use of multi-agency geospatial observations relevant for a common use case - without undergoing the search and discovery process in a less efficient, disparate path with each agency. NASA’s Earth Observing System Data and Information System (EOSDIS) and NOAA’s National Environmental Satellite, Data and Information Service (NESDIS) both support a wide range of Earth science disciplines’ research, operations, and applications activities. Presently, however, there are few examples of data discovery frameworks supporting an inquiry of both NASA’s and NOAA’s extensive archives of Earth observations that are equally suitable for a particular science scenario, regardless of the agency that “owns” the data. NASA and NOAA are collaborating on a data expedition platform for exploring fire weather using data products from both agencies. Users will be able to search, discover, and visualize NASA and NOAA products in one interface. Each agency will curate metadata for its respective datasets, providing for a rich search experience. The collaboration will pilot a shared search interface into these metadata datastores. Data products will be stored in the cloud in cloud-optimized format(s). These formats will allow for optimized data access and visualization to support the “data expedition”. Avenues for further development and application of this cloud-based, multi-agency data provisioning platform will also be discussed.

cloud-based technology↗

Data Integration Support for Data Served in the OPeNDAP and OGC Environments

NASA is coordinating a technology development project to construct a gateway between system components built upon the Open-source Project for a Network Data AcceSs Protocol (OPeNDAP) and those made available made available via interfaces specified by the Open Geospatial Consortium (OGC). This project is funded though the Advanced Collaborative Connections for Earth-Sun System Science (ACCESS) Program and is a NASA contribution to the Committee on Earth Satellites (CEOS) Working Group on Information Systems and Services (WGISS). The motivation for the project is the set of data integration needs that have been expressed by the Coordinated Enhanced Observing Period (CEOP), an international program that is addressing the study of the global water cycle. CEOP is assembling a large collection in situ and satellite data and mode1 results from a wide variety of sources covering 35 sites around the globe. The data are provided by systems based on either the OPeNDAP or OGC protocols but the research community desires access to the full range of data and associated services from a single client. This presentation will discuss the current status of the OPeNDAP/OGC Gateway Project. The project is building upon an early prototype that illustrated the feasibility of such a gateway and which was demonstrated to the CEOP science community. In its first year as an ACCESS project, the effort has been has focused on the design of the catalog and data services that will be provided by the gateway and the mappings between the metadata and services provided in the two environments.

McDonald, Kenneth R.↗