Search NASASearch

SEARCH · Search NASA

Results for “metadata evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Metadata Evaluation and Improvement: Evolving Analysis and Reporting

ESIP Community members create and manage a large collection of environmental datasets that span multiple decades, the entire globe, and many parts of the solar system. Metadata are critical for discovering, accessing, using and understanding these data effectively and ESIP community members have successfully created large collections of metadata describing these data. As part of the White House Big Earth Data Initiative (BEDI), ESDIS has developed a suite of tools for evaluating these metadata in native dialects with respect to recommendations from many organizations. We will describe those tools and demonstrate evolving techniques for sharing results with data providers.

metadata recommendations

Evaluating and Evolving Metadata in Multiple Dialects

Despite many long-term homogenization efforts, communities continue to develop focused metadata standards along with related recommendations and (typically) XML representations (aka dialects) for sharing metadata content. Different representations easily become obstacles to sharing information because each representation generally requires a set of tools and skills that are designed, built, and maintained specifically for that representation. In contrast, community recommendations are generally described, at least initially, at a more conceptual level and are more easily shared. For example, most communities agree that dataset titles should be included in metadata records although they write the titles in different ways.

metadata quality

Collection Evaluation and Evolution

We will review metadata evaluation tools and share results from our most recent CMR analysis. We will demonstrate results using Google spreadsheets and present new results in terms of number of records that include specific content. We will show evolution of UMM-compliance over time and also show results of comparing various CMR collections (NASA, non-NASA, and SciOps).

Evauation

pyQuARC: Open Source Library for Earth Observation Metadata Quality Assessment

Metadata quality is essential to effective data discovery and has become increasingly vital as more Earth Science data sets become available. The Common Metadata Repository (CMR) hosts metadata describing NASA’s Earth Observation data products, which are archived across 12 Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, conducts metadata quality assessments to ensure that these data products are discoverable, accessible, and usable. To achieve these goals, the ARC team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency. ARC uses a combination of manual and automated methods to assess these three components and identify areas of improvement; the team then collaborates with the DAACs to resolve any findings. To streamline this process, ARC is currently developing a host of scripts, known as pyQuARC, to automate metadata quality assessments as much as possible. pyQuARC is an open source library for Earth Observation Metadata Quality Assessment, and the tool utilizes ARC’s metadata quality assessment framework to make basic validation checks, pinpoint inconsistencies between dataset-level (i.e. collection) and file-level (i.e. granule) metadata, and identify opportunities for more descriptive and robust information. Since pyQuARC is also customizable, other users can make modifications as needed, and future metadata standards can also be implemented. Once pyQuARC is fully developed, it will support multiple schema types to serve the broader EOSDIS metadata community. This presentation will provide an overview of pyQuARC and its process of development while showcasing the tool’s valuable features and uses.

Jenny Wood

QuARC: Development of a Service to Enable FAIR-er Metadata

The ARC Project: The ARC Team located at NASA’s Marshall Space Flight Center conducts quality assessments of metadata records that catalog NASA’s collection of over 9,000 Earth observation data products, stored in a centralized database called the Common Metadata Repository (CMR). The ARC Team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency with the goal of making NASA’s data products more discoverable, accessible, and usable. ARC = Analysis and Review of the CMR

Earth Science Informatics

A Framework for Assessing Earth Observation Metadata Quality: Implications for Data Discovery and Open Science

The Common Metadata Repository (CMR) contains metadata records describing NASA’s collection of over 8,000 Earth observation data products. The Analysis and Review of CMR (ARC) Team at Marshall Space Flight Center assesses the quality of these metadata records. Metadata, rather than the data itself, is indexed for search in both discipline-specific datacenters and global or aggregated catalogs (such as Earth data Search), making it essential for determining whether a data product is appropriate for a given research question or application need. Since metadata connects users to data, it should be as accurate and complete as possible in addition to meeting minimum database requirements. The ARC team has developed a metadata quality framework by which to assess quality. The framework consists of a set of quality criteria that converge around the dimensions of correctness, completeness, and consistency, with the goal of improving the discoverability, accessibility, and usability of NASA’s Earth Observation data. The application of the framework has resulted in a measurable improvement in NASA’s metadata quality. Key aspects of the framework’s success are the ability to systematically evaluate metadata and provide actionable quality improvement recommendations. Lessons learned from the project will be shared along with implementation details which may be relevant to other science disciplines. By aiming to make data more discoverable and accessible to a broad user community, the ARC metadata quality framework helps contribute to NASA’s commitment to open science.

Jeanne Le Roux

Do-It-Yourself: A Special Library's Approach to Creating Dynamic Web Pages Using Commercial Off-The-Shelf Applications

Many librarians may feel that dynamic Web pages are out of their reach, financially and technically. Yet we are reminded in library and Web design literature that static home pages are a thing of the past. This paper describes how librarians at the Institute for Defense Analyses (IDA) library developed a database-driven, dynamic intranet site using commercial off-the-shelf applications. Administrative issues include surveying a library users group for interest and needs evaluation; outlining metadata elements; and, committing resources from managing time to populate the database and training in Microsoft FrontPage and Web-to-database design. Technical issues covered include Microsoft Access database fundamentals, lessons learned in the Web-to-database process (including setting up Database Source Names (DSNs), redesigning queries to accommodate the Web interface, and understanding Access 97 query language vs. Standard Query Language (SQL)). This paper also offers tips on editing Active Server Pages (ASP) scripting to create desired results. A how-to annotated resource list closes out the paper.

Relational databases

Documentation Resources on the ESIP Wiki

The ESIP community includes data providers and users that communicate with one another through datasets and metadata that describe them. Improving this communication depends on consistent high-quality metadata. The ESIP Documentation Cluster and the wiki play an important central role in facilitating this communication. We will describe and demonstrate sections of the wiki that provide information about metadata concept definitions, metadata recommendation, metadata dialects, and guidance pages. We will also describe and demonstrate the ISO Explorer, a tool that the community is developing to help metadata creators.

ESIP Documentation Cluster

WIS and WIGOS Metadata as the Foundation for a Sustainable Framework for Global Greenhouse Gas Watch Data Exchange

Metadata (data about data) is a critical component of data discovery, description, evaluation, documentation, and preservation. Developing and propagating metadata standards has been a longstanding area of activity in WMO and beyond. The WIS2 and WIGOS metadata models are being actively developed and maintained by dedicated task teams, established under the WMO Expert Team on Metadata. The metadata representations and vocabularies are governed by well-established processes within WMO. These standards are being used in a number of metadata/data exchange activities (e.g., WMO Information System 2.0 (WIS2), WIGOS (WMDR), Climate Data Management Systems (CMDS), etc.). It should also be noted that the application of the WIS2 and WIGOS standards fully support the WMO Unified Data Policy and open data policy as well as greatly enhance the value of observations by fostering data F.A.I.R.ness. Furthermore, the WMO metadata standards can serve as the foundation for a framework that will facilitate metadata mapping between the existing schemas used in well-established data centres, e.g., WMO WDCGG (World Data Centre for Greenhouse Gases) and NOAA ObsPack (Observation Package Data Products) and to automate metadata exchange between data centres as well as with WMO. These activities will play a central role in integrating measurements sponsored by various member countries and organizations to provide a more comprehensive characterization of the temporal and spatial distribution of the greenhouse gases. At the same time, this metadata exchange can lead to member countries and partner organizations improving their current metadata collection process for data discoverability, interoperability, and (re)usability. This presentation will describe metadata activities in the context of WIS2 and WIGOS and how they apply to GGGW data integration via metadata mapping and exchange.

Gao Chen

Commercial Smallsat Data Acquisition Program On-ramp #2 Airbus U.S. Synthetic Aperture Radar (SAR) Evaluation Report

In 2017, NASA’s Earth Science Division (ESD) launched the Private-Sector Small Constellation Satellite Data Product Pilot, now referred to as the Commercial Smallsat Data Acquisition (CSDA) program. The objective of CSDA is to identify, evaluate, and acquire commercial remote sensing data that support NASA’s Earth science research and application activities. The Pilot successfully concluded in early 2020, when CSDA transitioned into a sustained program with on-ramping opportunities for new vendors as the industry emerges with new candidates and capabilities. In October 2019, a Request for Information (RFI) seeking capability statements from parties interested in providing data from spaceborne platforms was released for the CSDA on-ramp #2 evaluations. To be responsive to the RFI, the commercial satellite constellations had to consist of three or more operating spacecraft actively collecting data in a non-geostationary orbit with full latitudinal coverage and be U.S. companies. Two vendors responded to the RFI and were evaluated by a committee composed of NASA ESD leadership, program managers, and scientists. Both vendors satisfied the RFI requirements and were asked to respond to a Request for Proposal (RFP). After review of the proposals, NASA entered into a Blanket Purchase Agreement (BPA) with Airbus Defense and Space GEO, Inc. (Airbus) U.S. in September 2021 and with BlackSky Geospatial Solutions, Inc. (BlackSky) in November 2021. In this report, CSDA provides an evaluation of the usefulness of data provided by the Airbus U.S. Synthetic Aperture Radar (SAR) satellite constellation, consisting of TerraSAR-X (launched in 2007), TanDEM-X (launched in 2010), and PAZ (launched in 2018), for advancing NASA’s Earth system science research and applications. The evaluation of the BlackSky commercial data will be provided in a separate report. To conduct the Airbus evaluation, NASA’s ESD augmented 13 existing research projects that could potentially benefit from, and had the expertise to evaluate, the commercial data being considered for longer-term purchase. Investigators from NASA’s Research and Analysis Program science focus areas and from NASA’s Applied Sciences Program elements participated in the evaluation. A summary of the research areas evaluated by the Principal Investigator (PI) teams is presented in Figure 3. CSDA also funded a dedicated activity to evaluate the satellite data quality (calibration and geolocation) independently by assessing the accuracy of data from Airbus. Evaluation activities were carried out by the selected PIs from December 7, 2022, to December 7, 2023. Delivery of datasets requested by the researchers began in January 2023. The vendors were evaluated on the accessibility of data, accuracy and completeness of metadata, and promptness and quality of user support services. Datasets purchased during the evaluation have been archived by NASA and will be made available to current and future government-funded researchers in accordance with the End User License Agreement (EULA). This synthesis report distills and integrates the findings of research reports commissioned by NASA for the Airbus evaluation. This report also includes recommendations that inform the way ahead for the program. The scientific results from the evaluations demonstrated that the commercial data from Airbus were able to advance NASA research and applications. However, the PIs encountered limitations that diminished the usefulness of the data due to the amount of effort that was required to access, preprocess, and analyze these data. One significant issue encountered was the limited spatial and temporal coverage of the data in the Airbus archive that could be used to conduct time series analyses or assessments over large spatial scales. Overall, however, the utility and the quality of the evaluated data outweighed the difficulties encountered, and NASA has concluded that the Airbus SAR data would complement NASA’s existing Earth observation capabilities and Airbus U.S. would qualify to participate in the sustained phase of the program.

Batuhan Osmanoglu

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert

GHRSST-14 DAS-TAG Report

The DAS-TAG provides the informatics and data management expertise in emerging information technologies for the GHRSST community. It provides expertise in data and metadata formats and standards, fosters improvements for GHRSST data curation, experiments with new data processing paradigms, and evaluates services and tools for data usage. It provides a forum for producer and distributor data management issues and coordination.

data processing

Community Involvement in Enhancing the Global Change Master Directory (GCMD) Controlled Vocabularies (Keywords)

NASA's Global Change Master Directory (GCMD) develops and expands a hierarchical set of controlled vocabularies (keywords) covering the Earth sciences and associated information (data centers, projects, platforms, instruments, etc.). The purpose of the keywords is to describe Earth science data and services in a consistent and comprehensive manner, allowing for the precise searching of metadata and subsequent retrieval of data and services. The keywords are accessible in a standardized SKOSRDFOWL representation and are used as an authoritative taxonomy, as a source for developing ontologies, and to search and access Earth Science data within online metadata catalogues. The keyword development approach involves: (1) receiving community suggestions, (2) triaging community suggestions, (3) evaluating the keywords against a set of criteria coordinated by the NASA ESDIS Standards Office, and (4) publication/notification of the keyword changes. This approach emphasizes community input, which helps ensure a high quality, normalized, and relevant keyword structure that will evolve with users changing needs. The Keyword Community Forum, which promotes a responsive, open, and transparent processes, is an area where users can discuss keyword topics and make suggestions for new keywords. The formalized approach could potentially be used as a model for keyword development.

governance

Commercial Smallsat Data Acquisition Program: Airbus U.S. Synthetic Aperture Radar Quality Assessment Summary

Quality assessment of the Airbus X-band Synthetic Aperture Radar (SAR) satellite products was conducted by the Commercial Smallsat Data Acquisition (CSDA) program’s radar subject matter experts, following the Joint NASA/ESA (European Space Agency) assessment draft guidelines. All three Airbus SAR spacecraft (TerraSAR-X, TanDEM-X, and PAZ) are based on the TerraSAR-X platform, and each have an active phased array antenna that is 4.8 x 0.7 m in the along-track and cross-track dimensions, respectively. TerraSAR-X and TanDEM-X are in a helical orbit, creating a bistatic imaging geometry, in addition to being capable of independent monostatic observations. The PAZ mission follows TerraSAR-X and TanDEM-X in the same 11-day orbit with a 5.5-day lag. TerraSAR-X and TanDEM-X are designed, developed, and operated through a Public-Private Partnership, while PAZ is a dual-use mission (civil and defense agencies), funded and owned by the Spanish Ministry of Defense and managed by Hisdesat (Hisdesat Servicios Estratégicos, S.A.), a Spanish private communications company. The assessment presented in this document is divided into two main parts: documentation review and the assessment of test datasets. The documentation review in sections 2.1 through 2.4 includes the assessment of the Airbus documentation provided to the CSDA evaluation team. The grading of these documents is given in columns 1-4 of the maturity matrix shown in section 1.1. Section 2.5 summarizes the evaluation performed by NASA using the data purchased through the CSDA program. The grading for this is given in the last column of the maturity matrix. Section 3 provides more detailed explanations on the methods and the results of the data analysis performed by NASA. Only the documents provided by Airbus for the evaluation were considered for the review. Additional documentation with more detailed description of the calibration and validation procedures may be available online but were not considered for this evaluation. The product information provided in the available documentation (RD-1, RD-2) and the product metadata together provided adequate information to work with the data. The product details in the metadata included the required information to work with the data in the common XML file format. Metrological traceability documentation was not provided to CSDA. All relevant characterization of the SAR system and data were provided, and the metadata include all relevant ancillary information. Documentation provided to CSDA included limited pre-flight and post-launch calibration information.

Batuhan Osmanoglu

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the “FAIRness” of NASA’s GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.

Evolving a NASA Digital Object Identifiers System with Community Engagement

To demonstrate how the ESDIS (Earth Science Data and Information System) DOI (Digital Object Identifier) system and its processes have evolved over these years based on the recommendations provided by the user community (whether the community members create and manage DOI information or use DOIs in the data citations). The user community is comprised of people with common interests and needs for data identifiers who are actively involved in the creation and usage process. Engagement describes the interactive context wherein the community provides information, evaluates the proposed processes, and provides guidance in the area of identifiers.

Identifiers

TOLNet’s FAIR Journey: Yesterday, Today, and Tomorrow

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America and contributed to several air quality focused field studies. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and ozone relevant research. As the broad scientific community has moved towards adopting FAIR Principles to make data more findable, accessible, interoperable, and (re)usable, the TOLNet team has been consistently making data more FAIR. This effort has many challenges, partially reflecting on the FAIR principles being domain agnostic while the implementation needs to be domain specific. The FAIR principles declare the dependence on the community standards, domain-relevant metadata, and rich metadata. This presentation uses the TOLNet data and data system as an example to explore the best practices to implement FAIR principle. Particularly, we will examine the metadata and the “richness” to support findability and usability as well as machine-to-machine actionability via API. Last year, as part of our FAIR journey, we launched the TOLNet website (https://tolnet.larc.nasa.gov/) and the API (https://tolnet.larc.nasa.gov/api/). Part of this process included extracting and cataloging metadata across the entire TOLNet mission timeframe. This enabled users to search through the mission by various metadata criteria, improving the findability and accessibility. And computers could connect directly to the TOLNet API to extract both metadata and data, providing a level of interoperability never present before for TOLNet data. On top of that, all new TOLNet data is now automatically validated using the API to ensure it complies with GEOMS standards, aiding in reusability. It takes both technology and scientists working together to make progress. The next step is to evaluate the current TOLNet offerings against NASA’s Practical Guide for Open, Free & FAIR NASA Earth Science Data Products (https://doi.org/10.5067/DOC/ESCO/ESDSWG-0002V1).

TOLNet