Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Application of a Dataset-Publication Knowledge Graph for Improving Earth Science Data Search

Finding a dataset at a NASA data center that is the best fit for the researcher’s application presents a challenge, not only for a novice user but for an experienced one, due to the data complexity and a multitude of choices of the existing data. Users often search for the data based on the application they are interested in, their research domain, phenomena, research topic, etc. As existing dataset metadata may not cover these search terms, the user may not obtain the most relevant results for their purpose. This problem was addressed by leveraging the content of the titles and abstracts of the research papers that utilize NASA datasets. For this, features from the paper titles and abstracts were extracted, and then a knowledge graph (KG) was used to link these features to the datasets used in that paper. The search for the datasets was tested by querying this knowledge graph through various terms extracted from Earth Science ontologies such as Semantic Web for Earth and Environment Technology (SWEET), and it was shown that this KG search outperforms the existing search that exclusively queries the dataset metadata.

Kristina Stoyanova↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.

Life Sciences data↗

Preserving NASA Historic and Current Mission Data and Adding Value to These for Future Researchers

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has been actively involved in many aspects of ensuring the long-term preservation of NASA earth science data and knowledge. This involves both the recovery and preservation of early NASA meteorological and other earth observation data, as well as preserving the more recent Earth Observation System (EOS) mission data sets which continue or have reached their end of lifetime. The GES DISC adds value to these preserved data by adding metadata and making the data available online to future researchers. The early NASA meteorological and earth observation data sets from the 1960s and 70s were originally archived on magnetic tapes, and visualizations of these data were preserved on 70-mm film. As these media have aged, their contents have been at risk of permanent loss. NASA has given the task of preserving these early data sets to the GES DISC and making these data sets easily available to the public. The data from these early missions are potentially useful to climate researchers as these are some of the only global measurements made at their time. These old data on magnetic tapes and film strips do not contain easily readable metadata, and so to add value the GES DISC has added digital metadata to them so that the data are searchable and findable. The GES DISC is also involved in preserving the data and knowledge from the EOS era missions. The GES DISC follows the guidelines developed for the preservation of data as specified in the NASA EOS Data and Information System (EOSDIS) Earth Science Data Preservation Content Specification (423-SPEC-001) document. To date, the GES DISC has consulted with the data science teams from the following missions: UARS, Earth Probe TOMS, Aura HIRDLS, and SORCE, in order to properly preserve their data and accompanying documentation. The GES DISC is also currently working with the EOS science teams from TRMM, AIRS, MLS, OMI and additional missions to ensure that the relevant documents and data sets are properly archived for future researchers. A standardized procedure for mission data preservation following 423-SPEC-001 makes preservation among the many NASA EOSDIS data centers uniform, so that these could be transitioned easily to a common EOSDIS preservation repository. This presentation will give an overview of the preservation and recovery of the old NASA historical data sets archived at the GES DISC, as well as the data and documentation preservation efforts of the EOS era missions.

James Johnson↗

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database↗

GES DISC Datalist Improves Earth Science Data Discoverability

At American Geophysical Union(AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a novel way to access data: Datalist. Currently, datalist is a collection of predefined data variables from one or more archived datasets, curated by our subject matter expert (SME). Our science support team has curated a predefined Hurricane Datalist and received very positive feedback from the user community. Datalist uses the same architecture our new website uses and have the same look and feel as other datasets on our web site. and also provides a one-stop shopping for data, metadata, citation, documentation, visualization and other available services. Since the last AGU Meeting, we have further developed a few new datalists corresponding to the Big Earth Data Initiative (BEDI) Societal Benefit Areas and A-Train data. We now have four datalists: Hurricane, Wind Energy, Greenhouse Gas and A-Train. We have also started working with our User Working Group members to create their favorite datalists and working with other DAAC to explore the possibility to include their products in our datalists that may also lead to a future of potential federated (cross-DAAC) datalists. Since our datalist prototype effort was a success, we are planning to make datalist operational. It's extremely important to have a common metadata model to support datalist, this will also be the foundation of federated datalist. We mapped our datalist metadata model to the unpublished UMM(Universal Metadata Model)-Var (Variable) (June version) and found that the UMM-var together with UMM-C (Collection) and possible UMM-S (Service) will meet our basic requirements. For example: Dataset shortname, and version are already specified in UMM-C, variable name, long name, units, dimensions are all specified in UMM-Var. UMM-Var also facilitates Science Keywords to allow tagging at variable level and Characteristics for optional variable characteristics. Measurements is useful for grouping of the variables and Set is promising to define datalist. And finally, the UMM-Service model to specify the available services for the variable will be very beneficial. In summary, UMM-Var, UMM-C and UMM-S are the basis of federated datalist and the development and deployment of datalist will contribute to the evolution of the UMM.

datalist↗

TOLNet’s FAIR Journey: Yesterday, Today, and Tomorrow

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America and contributed to several air quality focused field studies. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and ozone relevant research. As the broad scientific community has moved towards adopting FAIR Principles to make data more findable, accessible, interoperable, and (re)usable, the TOLNet team has been consistently making data more FAIR. This effort has many challenges, partially reflecting on the FAIR principles being domain agnostic while the implementation needs to be domain specific. The FAIR principles declare the dependence on the community standards, domain-relevant metadata, and rich metadata. This presentation uses the TOLNet data and data system as an example to explore the best practices to implement FAIR principle. Particularly, we will examine the metadata and the “richness” to support findability and usability as well as machine-to-machine actionability via API. Last year, as part of our FAIR journey, we launched the TOLNet website (https://tolnet.larc.nasa.gov/) and the API (https://tolnet.larc.nasa.gov/api/). Part of this process included extracting and cataloging metadata across the entire TOLNet mission timeframe. This enabled users to search through the mission by various metadata criteria, improving the findability and accessibility. And computers could connect directly to the TOLNet API to extract both metadata and data, providing a level of interoperability never present before for TOLNet data. On top of that, all new TOLNet data is now automatically validated using the API to ensure it complies with GEOMS standards, aiding in reusability. It takes both technology and scientists working together to make progress. The next step is to evaluate the current TOLNet offerings against NASA’s Practical Guide for Open, Free & FAIR NASA Earth Science Data Products (https://doi.org/10.5067/DOC/ESCO/ESDSWG-0002V1).

TOLNet↗

Open Science for Plants in Space: Improvements in NASA's Open Science Data Repository

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, elevated CO2, and many other abiotic stressors. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. Current OSDR standards include the ISA (Investigation-Study-Assay) experiment model, assay metadata configurations, and standardized terminology and ontologies. In 2024 OSDR will include a new suite of features for improved FAIR compliance including downloadable plant metadata templates, data submission tools and overall improved AI-readiness of plant datasets. AI/ML methods can be helpful tools to overcome the inherent challenges of space biology research (small sample size, sparse and heterogeneous data etc.). However these methods are built on an assumption of normalized and well-curated data. OSDR’s new curation tools will improve users ability to leverage ML and AI methods to model space biology data and better understand the complex effects of spaceflight on living systems across hierarchical biological levels. We look forward to sharing our advances with the spaceflight community.

FAIR↗

Recovery, Restoration and Archiving of Previously Lost Data and Metadata from the Apollo Lunar Surface Experiments Package (ALSEP)

The Apollo Lunar Surface Experiments Package (ALSEP) is the name used to collectively represent the geophysical instruments deployed on the lunar surface by the astronauts on Apollo 12, 14, 15, 16, and 17. These instruments were active from the times of their deployment (November 1969 – December 1972) to September 1977. During that time, fourteen types of experiments were conducted, and their data were transmitted to Earth. The experiment PIs processed them. At the conclusion of the experiments, some of these data were submitted to the NASA Space Science Data Coordinated Archive (NSSDCA) for archiving, while others were not. The raw instrument data received from the Moon prior to March 1976 were not archived, either. The unarchived data, resided on open-reel magnetic tapes, became lost in the decades since, along with much of the metadata (the information necessary/useful in properly processing/analyzing the data). This article retraces the history of the ALSEP data archiving efforts in the 1970s, the subsequent loss of the data tapes, and the search, recovery, and restoration of the lost data by contemporary researchers in the 21st century. In 2006, NSSDCA began reformatting some of the ALSEP data archived in the 1970s to conform with the current Planetary Data System (PDS). In 2010, 440 of the previously lost magnetic tapes containing the raw ALSEP data were recovered. From these tapes, the data were extracted, re-packaged for individual experiments, and, for those with sufficient metadata, processed into higher order data readily usable by researchers. All of these data products have been recently archived with either PDS or NSSDCA. These newly restored data fill a number of gaps in the previously existing archive of the ALSEP data. In addition, tens of thousands of pages of Apollo era documents have been optically scanned and compiled into an online searchable catalog. This article also describes the content, organization, and usage of the restored raw ALSEP data and metadata.

S Nagihara↗

Stewardship of NASA's Earth Science Data and Ensuring Long-Term Active Archives

Program, NASA has followed an open data policy, with non-discriminatory access to data with no period of exclusive access. NASA has well-established processes for assigning and or accepting datasets into one of 12 Distributed Active Archive Centers (DAACs) that are parts of EOSDIS. EOSDIS has been evolving through several information technology cycles, adapting to hardware and software changes in the commercial sector. NASA is responsible for maintaining Earth science data as long as users are interested in using them for research and applications, which is well beyond the life of the data gathering missions. For science data to remain useful over long periods of time, steps must be taken to preserve: (1) Data bits with no corruption, (2) Discoverability and access, (3) Readability, (4) Understandability, (5) Usability' and (6). Reproducibility of results. NASAs Earth Science data and Information System (ESDIS) Project, along with the 12 EOSDIS Distributed Active Archive Centers (DAACs), has made significant progress in each of these areas over the last decade, and continues to evolve its active archive capabilities. Particular attention is being paid in recent years to ensure that the datasets are published in an easily accessible and citable manner through a unified metadata model, a common metadata repository (CMR), a coherent view through the earthdata.gov website, and assignment of Digital Object Identifiers (DOI) with well-designed landing product information pages.

Data Management↗

JPL Project Information Management: A Continuum Back to the Future

This slide presentation reviews the practices and architecture that support information management at JPL. This practice has allowed concurrent use and reuse of information by primary and secondary users. The use of this practice is illustrated in the evolution of the Mars Rovers from the Mars Pathfinder to the development of the Mars Science Laboratory. The recognition of the importance of information management during all phases of a project life cycle has resulted in the design of an information system that includes metadata, has reduced the risk of information loss through the use of an in-process appraisal, shaping of project's appreciation for capturing and managing the information on one project for re-use by future projects as a natural outgrowth of the process. This process has also assisted in connection of geographically disbursed partners into a team through sharing information, common tools and collaboration.

Information Management↗

EED-2 Contract Final Report

This EED-2 contract provides for development and sustaining engineering of software and hardware systems that provide science data management for the ESDIS Project. A major activity under this contract will be for evolution and development engineering of the EOSDIS Core System (ECS), Earthdata platform, Common Metadata Repository (CMR), NASA-Compliant General Application Platform (NGAP) cloud environment, and other EOSDIS elements that provide the common capabilities and infrastructure of EOSDIS.

EOSDIS↗

Stewardship Best Practices for Improved Discovery and Reuse of Heterogeneous and Cross-Disciplinary Earth System Data

Some of the Earth system data products such as those from NASA airborne and field investigations (a.k.a. campaigns), are highly heterogeneous and cross-disciplinary, making the data extremely challenging to manage. For example, airborne and field campaign measurements tend to be sporadic over a period of time, with large gaps. Data products generated are of various processing levels and utilized for a wide range of inter- and cross-disciplinary research and applications. Data and derived products have been historically stored in a variety of domain-specific standard (and some non-standard) formats and in various locations such as NASA Distributed Active Archive Centers (DAACs), NASA airborne science facilities, field archives, or even individual scientists’ computer hard drives. As a result, airborne and field campaign data products have often been managed and represented differently, making it onerous for data users to find, access, and utilize campaign data. Some difficulties in discovering and accessing the campaign data originate from the incomplete data product and contextual metadata that may contain details relevant to the campaign (e.g. campaign acronym and instrument deployment locations), but tend to lack other significant information needed to understand conditions surrounding the data. Such details can be burdensome to locate after the conclusion of a campaign. Utilizing consistent terminology, essential for improved discovery and reuse, is also challenging due to the variety of involved disciplines. To help address the aforementioned challenges faced by many repositories and data managers handling airborne and field data, this presentation will describe stewardship practices developed by the Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) under the NASA’s Earth Science Data systems (ESDS) Program.

best practices↗

Community Coordinated Modeling Center Support of Science Needs for Integrated Data Environment

Space science models are essential component of integrated data environment. Space science models are indispensable tools to facilitate effective use of wide variety of distributed scientific sources and to place multi-point local measurements into global context. The Community Coordinated Modeling Center (CCMC) hosts a set of state-of-the- art space science models ranging from the solar atmosphere to the Earth's upper atmosphere. The majority of models residing at CCMC are comprehensive computationally intensive physics-based models. To allow the models to be driven by data relevant to particular events, the CCMC developed an online data file generation tool that automatically downloads data from data providers and transforms them to required format. CCMC provides a tailored web-based visualization interface for the model output, as well as the capability to download simulations output in portable standard format with comprehensive metadata and user-friendly model output analysis library of routines that can be called from any C supporting language. CCMC is developing data interpolation tools that enable to present model output in the same format as observations. CCMC invite community comments and suggestions to better address science needs for the integrated data environment.

Kuznetsova, M. M.↗

Earth Science Markup Language: Transitioning From Design to Application

The primary objective of the proposed Earth Science Markup Language (ESML) research is to transition from design to application. The resulting schema and prototype software will foster community acceptance for the "define once, use anywhere" concept central to ESML. Supporting goals include: 1. Refinement of the ESML schema and software libraries in cooperation with the user community. 2. Application of the ESML schema and software libraries to a variety of Earth science data sets and analysis tools. 3. Development of supporting prototype software for enhanced ease of use. 4. Cooperation with standards bodies in order to assure ESML is aligned with related metadata standards as appropriate. 5. Widespread publication of the ESML approach, schema, and software.

Moe, Karen↗

ESML for Earth Science Data Sets and Analysis

The primary objective of this research project was to transition ESML from design to application. The resulting schema and prototype software will foster community acceptance for the Define once, use anywhere concept central to ESML. Supporting goals include: 1) Refinement of the ESML schema and software libraries in cooperation with the user community; 2) Application of the ESML schema and software to a variety of Earth science data sets and analysis tools; 3) Development of supporting prototype software for enhanced ease of use; 4) Cooperation with standards bodies in order to assure ESML is aligned with related metadata standards as appropriate; and 5) Widespread publication of the ESML approach, schema, and software.

Graves, Sara↗

Report on the Global Data Assembly Center (GDAC) to the 12th GHRSST Science Team Meeting

In 2010/2011 the Global Data Assembly Center (GDAC) at NASA's Physical Oceanography Distributed Active Archive Center (PO.DAAC) continued its role as the primary clearinghouse and access node for operational Group for High Resolution Sea Surface Temperature (GHRSST) datastreams, as well as its collaborative role with the NOAA Long Term Stewardship and Reanalysis Facility (LTSRF) for archiving. Here we report on our data management activities and infrastructure improvements since the last science team meeting in June 2010.These include the implementation of all GHRSST datastreams in the new PO.DAAC Data Management and Archive System (DMAS) for more reliable and timely data access. GHRSST dataset metadata are now stored in a new database that has made the maintenance and quality improvement of metadata fields more straightforward. A content management system for a revised suite of PO.DAAC web pages allows dynamic access to a subset of these metadata fields for enhanced dataset description as well as discovery through a faceted search mechanism from the perspective of the user. From the discovery and metadata standpoint the GDAC has also implemented the NASA version of the OpenSearch protocol for searching for GHRSST granules and developed a web service to generate ISO 19115-2 compliant metadata records. Furthermore, the GDAC has continued to implement a new suite of tools and services for GHRSST datastreams including a Level 2 subsetter known as Dataminer, a revised POET Level 3/4 subsetter and visualization tool, a Google Earth interface to selected daily global Level 2 and Level 4 data, and experimented with a THREDDS catalog of GHRSST data collections. Finally we will summarize the expanding user and data statistics, and other metrics that we have collected over the last year demonstrating the broad user community and applications that the GHRSST project continues to serve via the GDAC distribution mechanisms. This report also serves by extension to summarize the activities of the GHRSST Data Assembly and Systems Technical Advisory Group (DAS-TAG).

sea surface temperature (SST)↗

Taming Big Data Variety in the Earth Observing System Data and Information System

Although the volume of the remote sensing data managed by the Earth Observing System Data and Information System is formidable, an oft-overlooked challenge is the variety of data. The diversity in satellite instruments, science disciplines and user communities drives cost as much or more as the data volume. Several strategies are used to tame this variety: data allocation to distinct centers of expertise; a common metadata repository for discovery, data format standards and conventions; and services that further abstract the variations in data.

Information Systems↗