Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata improvement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Questions/Issues to be Discussed at the Snow/Ice Workshop

How soon after acquisition will you need the snow/ice maps? For the composite maps, which do you prefer, a composite of 7 days, 10 days, other, and why? What would be the most useful MODIS at-launch and post-launch snow and ice products? Specifically what would you use the products for? What metadata should be included with the data products? For example, quality control data are metadata. Image i.d.# and lat/long are also metadata. What improvements can you suggest to the snow and ice products as currently planned?

Hall, Dorothy K.↗

Sally Ride EarthKAM - Automated Image Geo-Referencing Using Google Earth Web Plug-In

Sally Ride EarthKAM is an educational program funded by NASA that aims to provide the public the ability to picture Earth from the perspective of the International Space Station (ISS). A computer-controlled camera is mounted on the ISS in a nadir-pointing window; however, timing limitations in the system cause inaccurate positional metadata. Manually correcting images within an orbit allows the positional metadata to be improved using mathematical regressions. The manual correction process is time-consuming and thus, unfeasible for a large number of images. The standard Google Earth program allows for the importing of KML (keyhole markup language) files that previously were created. These KML file-based overlays could then be manually manipulated as image overlays, saved, and then uploaded to the project server where they are parsed and the metadata in the database is updated. The new interface eliminates the need to save, download, open, re-save, and upload the KML files. Everything is processed on the Web, and all manipulations go directly into the database. Administrators also have the control to discard any single correction that was made and validate a correction. This program streamlines a process that previously required several critical steps and was probably too complex for the average user to complete successfully. The new process is theoretically simple enough for members of the public to make use of and contribute to the success of the Sally Ride EarthKAM project. Using the Google Earth Web plug-in, EarthKAM images, and associated metadata, this software allows users to interactively manipulate an EarthKAM image overlay, and update and improve the associated metadata. The Web interface uses the Google Earth JavaScript API along with PHP-PostgreSQL to present the user the same interface capabilities without leaving the Web. The simpler graphical user interface will allow the public to participate directly and meaningfully with EarthKAM. The use of similar techniques is being investigated to place ground-based observations in a Google Mars environment, allowing the MSL (Mars Science Laboratory) Science Team a means to visualize the rover and its environment.

Andres, Paul M.↗

WIS and WIGOS Metadata as the Foundation for a Sustainable Framework for Global Greenhouse Gas Watch Data Exchange

Metadata (data about data) is a critical component of data discovery, description, evaluation, documentation, and preservation. Developing and propagating metadata standards has been a longstanding area of activity in WMO and beyond. The WIS2 and WIGOS metadata models are being actively developed and maintained by dedicated task teams, established under the WMO Expert Team on Metadata. The metadata representations and vocabularies are governed by well-established processes within WMO. These standards are being used in a number of metadata/data exchange activities (e.g., WMO Information System 2.0 (WIS2), WIGOS (WMDR), Climate Data Management Systems (CMDS), etc.). It should also be noted that the application of the WIS2 and WIGOS standards fully support the WMO Unified Data Policy and open data policy as well as greatly enhance the value of observations by fostering data F.A.I.R.ness. Furthermore, the WMO metadata standards can serve as the foundation for a framework that will facilitate metadata mapping between the existing schemas used in well-established data centres, e.g., WMO WDCGG (World Data Centre for Greenhouse Gases) and NOAA ObsPack (Observation Package Data Products) and to automate metadata exchange between data centres as well as with WMO. These activities will play a central role in integrating measurements sponsored by various member countries and organizations to provide a more comprehensive characterization of the temporal and spatial distribution of the greenhouse gases. At the same time, this metadata exchange can lead to member countries and partner organizations improving their current metadata collection process for data discoverability, interoperability, and (re)usability. This presentation will describe metadata activities in the context of WIS2 and WIGOS and how they apply to GGGW data integration via metadata mapping and exchange.

Gao Chen↗

Optimizing Metadata Exchange: Leveraging DAOS for ADIOS Metadata I/O

In HPC I/O middleware like the Adaptable I/O System (ADIOS) often mediates data transfers between applications. The metadata I/O generated by such systems often presents significant scaling and performance limitations. This work seeks improvement opportunities for metadata I/O by leveraging the DAOS storage systems, a recent storage system solution deployed on high-end systems such as the Aurora supercomputer. We investigate the tradeoffs and the design space for integrating I/O engines for the ADIOS middleware based on the different storage mechanisms supported by DAOS. We present a new DAOS-Array-ChunkSize-aligned engine which provides up to 2.3× improved performance than when using the existing DAOS-POSIX interface, without requiring any application modifications.

Venkatesh, Ranjan Sarpangala↗

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation↗

A Framework for Assessing Earth Observation Metadata Quality: Implications for Data Discovery and Open Science

The Common Metadata Repository (CMR) contains metadata records describing NASA’s collection of over 8,000 Earth observation data products. The Analysis and Review of CMR (ARC) Team at Marshall Space Flight Center assesses the quality of these metadata records. Metadata, rather than the data itself, is indexed for search in both discipline-specific datacenters and global or aggregated catalogs (such as Earth data Search), making it essential for determining whether a data product is appropriate for a given research question or application need. Since metadata connects users to data, it should be as accurate and complete as possible in addition to meeting minimum database requirements. The ARC team has developed a metadata quality framework by which to assess quality. The framework consists of a set of quality criteria that converge around the dimensions of correctness, completeness, and consistency, with the goal of improving the discoverability, accessibility, and usability of NASA’s Earth Observation data. The application of the framework has resulted in a measurable improvement in NASA’s metadata quality. Key aspects of the framework’s success are the ability to systematically evaluate metadata and provide actionable quality improvement recommendations. Lessons learned from the project will be shared along with implementation details which may be relevant to other science disciplines. By aiming to make data more discoverable and accessible to a broad user community, the ARC metadata quality framework helps contribute to NASA’s commitment to open science.

Jeanne Le Roux↗

Reusing Data and Metadata to Create New Metadata Through Machine-Learning & Other Programmatic Methods

Recent improvements in natural language processing (NLP) enable metadata to be created programmatically from reused original metadata or even the dataset itself. Transfer-learning applied to NLP has greatly improved performance and reduced training data requirements. In this talk, we’ll compare machine-generated metadata to human-generated metadata and discuss characteristics of metadata and data archives that affect suitability for machine-learning reuse of metadata. Where as human-generated metadata is often populated once, populated from the perspective of data supplier, populated by many individuals with different words for the same thing, and limited in length, machine-generated metadata can be updated any number of times, generated from the perspective of any user, constrained to a standardized set of terms that can be evolved over time, and be any length required. Machine-learning generated metadata offers benefits but also additional needs in terms of version control, process transparency, human-computer interaction, and IT requirements. As a successful example, we’ll discuss how a dataset of abstracts and associated human-tagged keywords from a standardized list of several thousand keywords were used to create a machine-learning model that predicted keyword metadata for open-source code projects on code.nasa.gov. We’ll also discuss a less successful example from data.nasa.gov to show how data archive architecture and characteristics of initial metadata can be strong controls on how easy it is to leverage programmatic methods to reuse metadata to create additional metadata.

Gosses, Justin↗

TomoPyUI : a user-friendly tool for rapid tomography alignment and reconstruction

The management and processing of synchrotron and neutron computed tomography data can be a complex, labor-intensive and unstructured process. Users devote substantial time to both manually processing their data ( i.e. organizing data/metadata, applying image filters etc. ) and waiting for the computation of iterative alignment and reconstruction algorithms to finish. In this work, we present a solution to these problems: TomoPyUI , a user interface for the well known tomography data processing package TomoPy . This highly visual Python software package guides the user through the tomography processing pipeline from data import, preprocessing, alignment and finally to 3D volume reconstruction. The TomoPyUI systematic intermediate data and metadata storage system improves organization, and the inspection and manipulation tools (built within the application) help to avoid interrupted workflows. Notably, TomoPyUI operates entirely within a Jupyter environment. Herein, we provide a summary of these key features of TomoPyUI , along with an overview of the tomography processing pipeline, a discussion of the landscape of existing tomography processing software and the purpose of TomoPyUI , and a demonstration of its capabilities for real tomography data collected at SSRL beamline 6-2c.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Astronaut Photography of the Earth: A Long-Term Dataset for Earth Systems Research, Applications, and Education

The NASA Earth observations dataset obtained by humans in orbit using handheld film and digital cameras is freely accessible to the global community through the online searchable database at https://eol.jsc.nasa.gov, and offers a useful compliment to traditional ground-commanded sensor data. The dataset includes imagery from the NASA Mercury (1961) through present-day International Space Station (ISS) programs, and currently totals over 2.6 million individual frames. Geographic coverage of the dataset includes land and oceans areas between approximately 52 degrees North and South latitudes, but is spatially and temporally discontinuous. The photographic dataset includes some significant impediments for immediate research, applied, and educational use: commercial RGB films and camera systems with overlapping bandpasses; use of different focal length lenses, unconstrained look angles, and variable spacecraft altitudes; and no native geolocation information. Such factors led to this dataset being underutilized by the community but recent advances in automated and semi-automated image geolocation, image feature classification, and web-based services are adding new value to the astronaut-acquired imagery. A coupled ground software and on-orbit hardware system for the ISS is in development for planned deployment in mid-2017; this system will capture camera pose information for each astronaut photograph to allow automated, full georegistration of the data. The ground system component of the system is currently in use to fully georeference imagery collected in response to International Disaster Charter activations, and the auto-registration procedures are being applied to the extensive historical database of imagery to add value for research and educational purposes. In parallel, machine learning techniques are being applied to automate feature identification and classification throughout the dataset, in order to build descriptive metadata that will improve search capabilities. It is expected that these value additions will increase interest and use of the dataset by the global community.

Stefanov, William L.↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

Data Citation Explorer (DCE) v1.0

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data whether or not provenance was clearly indicated. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

Parker, Charles↗

ICARTT File Format Enhancements: Supporting FAIRness and Data Discovery of Suborbital Campaign Data

Suborbital campaigns aim to accomplish a wide variety of goals and can include a variety of platforms, instruments, and parameters measured. In 2004, the ICARTT (International Consortium for Atmospheric Research on Transport and Transformation) standards were developed to fulfill data management needs for the ICARTT campaign. The ICARTT file format is text-based and composed of a header with important data description information and the data section. Built on the NASA Ames and GTE data formats, the ICARTT format was created to facilitate data exchange and promote collaborations among the science teams for achieving the ICARTT campaign goals. Due to its success and adaptation for use in many other field campaigns, the ICARTT file format became a NASA standard in 2010 and was amended in January 2017. These changes provided many enhancements, including the requirement for variable standard names. Primarily designed for airborne field studies, ICARTT has been further utilized for ground-based studies. NASA has made a commitment to build an inclusive open science community over the next decade. Open-source science strives to make publicly funded scientific research transparent, inclusive, accessible, and reproducible. The ICARTT format can host metadata that is critical for proper use of the data, particularly for in-situ measurements, and can enhance data discovery and accessibility. However, the required fields are often free text, meaning that the information is human readable, but not machine interpretable. Furthermore, the amount and type of information provided can vary significantly between principal investigators and campaigns. To support FAIR principles and interoperability, enhancements to the ICARTT standards are recommended. Possible recommendations include potential use of controlled and consistent vocabulary for variable standard name and certain common metadata elements; standardizing timestamps for easier data comparisons and analysis; and providing guidance on variable measurement units and how they are reported. Enhancing ICARTT metadata can further streamline the process to make suborbital data more readily available to the data user and improve variable-level metadata. Providing more variable-level metadata can enhance data searching and discovery, supporting NASA’s Open-Source Science Initiative (OSSI).

Megan Buzanowicz↗

Making Metadata Better with CMR and MMT

Ensuring complete, consistent and high quality metadata is a challenge for metadata providers and curators. The CMR and MMT systems provide providers and curators options to build in metadata quality from the start and also assess and improve the quality of already existing metadata.

Quality↗

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert↗

Mercury Toolset for Spatiotemporal Metadata

Mercury (http://mercury.ornl.gov) is a set of tools for federated harvesting, searching, and retrieving metadata, particularly spatiotemporal metadata. Version 3.0 of the Mercury toolset provides orders of magnitude improvements in search speed, support for additional metadata formats, integration with Google Maps for spatial queries, facetted type search, support for RSS (Really Simple Syndication) delivery of search results, and enhanced customization to meet the needs of the multiple projects that use Mercury. It provides a single portal to very quickly search for data and information contained in disparate data management systems, each of which may use different metadata formats. Mercury harvests metadata and key data from contributing project servers distributed around the world and builds a centralized index. The search interfaces then allow the users to perform a variety of fielded, spatial, and temporal searches across these metadata sources. This centralized repository of metadata with distributed data sources provides extremely fast search results to the user, while allowing data providers to advertise the availability of their data and maintain complete control and ownership of that data. Mercury periodically (typically daily) harvests metadata sources through a collection of interfaces and re-indexes these metadata to provide extremely rapid search capabilities, even over collections with tens of millions of metadata records. A number of both graphical and application interfaces have been constructed within Mercury, to enable both human users and other computer programs to perform queries. Mercury was also designed to support multiple different projects, so that the particular fields that can be queried and used with search filters are easy to configure for each different project.

Wilson, Bruce E.↗

A User-Focused Renovation of CERES Metadata

Production software and public data products for Clouds and the Earth’s Radiant Energy System (CERES) continue to evolve as the project extends its climate data record. The data management team for CERES is currently undertaking major renovations of both code and data products, the latter of which is, of course, in service of improving user experience. A major mode of CERES’ data product improvement is in renovating products’ metadata. Metadata standards have evolved since CERES began producing its data products in 2000. In its twentieth year, CERES essentially asked the question: how would the project design its data products if it could start all over again? With forthcoming editions, this rebirth will be realized. CERES has redesigned its metadata standards to best position itself for data discoverability. The project has used the latest standards being developed in NASA’s Earth Science Data and Information Systems (ESDIS) Project’s Unified Metadata Model (UMM) documentation; collaborated with the Atmospheric Science Data Center (ASDC) to ensure compliance with Common Metadata Repository compatibility, and continued compliance with Climate and Forecast (CF) Conventions. In doing so, the team created its own, internal document for proper metadata creation and metadata verification software that is deployed prior to all code deliveries. This presentation will discuss this redesign process, as well as needs met and those that are still outstanding in the search for an improved user experience with CERES data products.

Kathleen Dejwakh↗

Report on the Global Data Assembly Center (GDAC) to the 12th GHRSST Science Team Meeting

In 2010/2011 the Global Data Assembly Center (GDAC) at NASA's Physical Oceanography Distributed Active Archive Center (PO.DAAC) continued its role as the primary clearinghouse and access node for operational Group for High Resolution Sea Surface Temperature (GHRSST) datastreams, as well as its collaborative role with the NOAA Long Term Stewardship and Reanalysis Facility (LTSRF) for archiving. Here we report on our data management activities and infrastructure improvements since the last science team meeting in June 2010.These include the implementation of all GHRSST datastreams in the new PO.DAAC Data Management and Archive System (DMAS) for more reliable and timely data access. GHRSST dataset metadata are now stored in a new database that has made the maintenance and quality improvement of metadata fields more straightforward. A content management system for a revised suite of PO.DAAC web pages allows dynamic access to a subset of these metadata fields for enhanced dataset description as well as discovery through a faceted search mechanism from the perspective of the user. From the discovery and metadata standpoint the GDAC has also implemented the NASA version of the OpenSearch protocol for searching for GHRSST granules and developed a web service to generate ISO 19115-2 compliant metadata records. Furthermore, the GDAC has continued to implement a new suite of tools and services for GHRSST datastreams including a Level 2 subsetter known as Dataminer, a revised POET Level 3/4 subsetter and visualization tool, a Google Earth interface to selected daily global Level 2 and Level 4 data, and experimented with a THREDDS catalog of GHRSST data collections. Finally we will summarize the expanding user and data statistics, and other metrics that we have collected over the last year demonstrating the broad user community and applications that the GHRSST project continues to serve via the GDAC distribution mechanisms. This report also serves by extension to summarize the activities of the GHRSST Data Assembly and Systems Technical Advisory Group (DAS-TAG).

sea surface temperature (SST)↗