Search NASA⌕ Search

SEARCH · Search NASA

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Artificial Intelligence Enhancements to Imagery for Space Operations

Philosophy classes still ponder the question asked by Dr. George Berkely, an Anglican Bishop and philosopher in the 1600’s-- “If a tree falls in a forest and no one is around to hear it, does it make a sound?” With that in mind, I ask the following—If a still image or motion imagery from a space mission cannot be found during a search, does it exist? Since the beginning of spaceflight, imagery has been a key form of data collected. Whether for mere curiosity (what does Earth look like from Space?), or for operational reasons (did the solar panel deploy?), or for engineering purposes (what was that object that floated away from the spacecraft?), imagery has been included in space missions. To be useful, though, the image or motion imagery must be accessible and accessed when needed. During the analog era, that typically meant captions and numbers associated with the physical media. With “born digital” imagery, it is possible to add metadata to the image data file. This metadata might include the date and time of capture, mission, camera, exposure data, and similar data fields. Many modern cameras embed some basic metadata into the image file at the moment of capture. The reality, though, is even with today’s born-digital enhancements with embedded metadata at the time of capture, reviewing and cataloging still and motion imagery is very labor intensive. Humans review the imagery for sensitive content (privacy concerns, imagery containing proprietary data/subject matter), and to identify imagery containing crew members or imagery that should be reviewed for engineering or scientific reasons. All this review and manual data entry is very time-consuming. Many improvements in Artificial Intelligence (AI), Machine Learning, and processing power now make it possible to identify persons, objects, motion, color, audio with sensitive content, and other details after or while the imagery is captured.

Rodney Grubbs↗

Data Sharing in Radiobiology; Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally „Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Data Sharing in Radiation Biology: Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally "Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Enabling Cloud Services and Enhanced Data Discovery With Earthdata-Varinfo

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. Each of these collections can contain hundreds of variables that describe each measurement within the dataset, therefore an automated method for generating UMM-Var records is necessary. The Unified Metadata Model for Variables (UMM-Var) provides a framework for variable metadata records in NASA’s Common Metadata Repository (CMR). The Python tool, earthdata-varinfo, was developed to solve this problem of automating the curation of UMM-Var records. Given either a collection DMR file or a netCDF-4 file, earthdata-varinfo can scrape variable metadata and return a CMR compliant UMM-Var record. Earthdata-varinfo can generate thousands of UMM-Var records in a matter of seconds, thus enabling subsetting capabilities and enhancing data discovery.

Eni Awowale↗

Extending CF Conventions to Enhance Data FAIRness for Atmospheric Composition Observations

The Hierarchical Data Format (HDF) and Network Common Data Form (NetCDF) are data file formats created to aid users in the creation or use of scientific data. These file formats are useful for handling large data volumes and hosting extensive metadata as global, group, or variable attributes and are popular with the modeling community. HDF and NetCDF files are widely used with atmospheric remote sensing data and have been used to support measurements from numerous field campaigns, from satellite to aircraft or ground and mobile based measurements. The files from airborne field studies, however, vary greatly in terms of the file structure and the amount and content of their metadata. Information relevant to the file that can be useful to the user such as the data producer, location where data was taken, variable descriptions, or information about the instrument might not be included in the file. Recently, the Measurements of Aerosols, Clouds, and their Interactions for Earth System Models (MACIE) group started a grassroots effort to develop a CF-based template for the HDF and NetCDF files for field studies, with the aim of making the data products more interoperable and usable. This template seeks to make the files more compliant to Climate and Forecast (CF) metadata conventions and to standardize the file structure and the global and variable attributes. The template would help to ensure that HDF and NetCDF files contain adequate metadata to better support their use for research, e.g., the modeling community, and to enhance the usability and interoperability of data for research communities at large. The draft template has been applied to recent field studies for various instruments and their merge files in support of the Atmosphere Observing System (AOS) project. The details of the revised template are to be presented, as well as examples of the implementation of these requirements for merge files and lidar observation data files and issues revealed during the implementation process.

Sean Leavor↗

Interoperability and Other Aspects of Guiding Data Producers for the Benefit of End Users

The purpose of this paper is to discuss how the Climate and Forecast (CF) Metadata Conventions and netCDF standard have influenced the recommendations and guidance provided to producers of data products based on NASA’s Earth observations. It has been long-recognized that interoperable datasets and use of standards and conventions are beneficial to the users of these datasets, especially those who make use of multiple datasets for their research and applications. The Dataset Interoperability Working Group (DIWG), one of NASA’s Earth Science Data System Working Groups (ESDSWGs), was established in 2013, and has developed and published many recommendations. The Data Product Development Guide (DPDG) Working Group, established in 2018 as another of the ESDSWGs, has published a DPDG for Data Producers and a Quick Start Guide, incorporating guidance from many sources, including the recommendations from the DIWG. The DPDG includes recommendations regarding data formats (prominently netCDF-4) and metadata based primarily on the CF Metadata Conventions and the Attribute Convention for Data Discovery (ACDD). In early 2023, it was decided that the Resource Center for Data Producers (RCDP) Working Group be established as another ESDSWG, with the goals of providing all the information relevant and helpful for data producers via an easily accessible website, and of recommending how the DPDG and QSG could be maintained as living documents, given the rapidly changing technologies, and the need for incorporating the experience and feedback from the users of these documents.

Data product development↗

Standard formatted data units-control authority operations

The purpose of this document is to illustrate a Control Authority's (CA) possible operation. The document is an interpretation and expansion of the concept found in the CA Procedures Recommendation. The CA is described in terms of the functions it performs for the management and control of data descriptions (metadata). Functions pertaining to the organization of Member Agency Control Authority Offices (MACAOs) (e.g., creating and disbanding) are not discussed. The document also provides an illustrative operational view of a CA through scenarios describing interaction between those roles involved in collecting, controlling, and accessing registered metadata. The roles interacting with the CA are identified by their actions in requesting and responding to requests for metadata, and by the type of information exchanged. The scenarios and examples presented in this document are illustrative only. They represent possible interactions supported by either a manual or automated system. These scenarios identify requirements for an automated system. These requirements are expressed by identifying the information to be exchanged and the services that may be provided by a CA for that exchange.

Source record↗

Technical Report Interchange Through Synchronized OAI Caches

The Technical Report Interchange project is a cooperative experimental effort between NASA Langley Research Center, Los Alamos National Laboratory, Air Force Research Laboratory, Sandia National Laboratory and Old Dominion University to allow for the integration of technical reports. This is accomplished using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) and having each site cache the metadata from the other participating sites. Each site also implements additional software to ingest the OAI-PMH harvested metadata into their native digital library (DL). This allows the users at each site to see an increased technical report collection through the familiar DL interfaces and tale advantage of whatever valued added are provided by the native DL.

Liu, Xiaming↗

ISAIA: Interoperable Systems for Archival Information Access

The ISAIA project was originally proposed in 1999 as a successor to the informal AstroBrowse project. AstroBrowse, which provided a data location service for astronomical archives and catalogs, was a first step toward data system integration and interoperability. The goals of ISAIA were ambitious: '...To develop an interdisciplinary data location and integration service for space science. Building upon existing data services and communications protocols, this service will allow users to transparently query hundreds or thousands of WWW-based resources (catalogs, data, computational resources, bibliographic references, etc.) from a single interface. The service will collect responses from various resources and integrate them in a seamless fashion for display and manipulation by the user.' Funding was approved only for a one-year pilot study, a decision that in retrospect was wise given the rapid changes in information technology in the past few years and the emergence of the Virtual Observatory initiatives in the US and worldwide. Indeed, the ISAIA pilot study was influential in shaping the science goals, system design, metadata standards, and technology choices for the virtual observatory. The ISAIA pilot project also helped to cement working relationships among the NASA data centers, US ground-based observatories, and international data centers. The ISAIA project was formed as a collaborative effort between thirteen institutions that provided data to astronomers, space physicists, and planetary scientists. Among the fruits we ultimately hoped would come from this project would be a central site on the Web that any space scientist could use to efficiently locate existing data relevant to a particular scientific question. Furthermore, we hoped that the needed technology would be general enough to allow smaller, more-focused community within space science could use the same technologies and standards to provide more specialized services. A major challenge to searching for data across a broad community is that information that describe some data products are either not relevant to other data or not applicable in the same way. Some previous metadata standard development efforts (e.g., in the earth science and library communities) have produced standards that are very large and difficult to support. To address this problem, we studied how a standard may be divided into separable pieces. Data providers that wish to participate in interoperable searches can support only those parts of the standard that are relevant to them. We prototyped a top-level metadata standard that was small and applicable to all space science data.

Hanisch, Robert J.↗

Questions/Issues to be Discussed at the Snow/Ice Workshop

How soon after acquisition will you need the snow/ice maps? For the composite maps, which do you prefer, a composite of 7 days, 10 days, other, and why? What would be the most useful MODIS at-launch and post-launch snow and ice products? Specifically what would you use the products for? What metadata should be included with the data products? For example, quality control data are metadata. Image i.d.# and lat/long are also metadata. What improvements can you suggest to the snow and ice products as currently planned?

Hall, Dorothy K.↗

OAI and NASA's Scientific and Technical Information

The Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) is an evolving protocol and philosophy regarding interoperability for digital libraries (DLs). Previously, "distributed searching" models were popular for DL interoperability. However, experience has shown distributed searching systems across large numbers of DLs to be difficult to maintain in an Internet environment. The OAI-PMH is a move away from distributed searching, focusing on the arguably simpler model of "metadata harvesting". We detail NASA s involvement in defining and testing the OAI-PMH and experience to date with adapting existing NASA distributed searching DLs (such as the NASA Technical Report Server) to use the OAI-PMH and metadata harvesting. We discuss some of the entirely new DL projects that the OAI-PMH has made possible, such as the Technical Report Interchange project. We explain the strategic importance of the OAI-PMH to the mission of NASA s Scientific and Technical Information Program.

Nelson, Michael L.↗

ECHO Status for International Partners

The EOS Clearinghouse (ECHO) is a clearinghouse of spatial and temporal metadata, inclusive of NASA's Distributed Active Archive Center (DAAC) data holdings, that enables the science community to more easily exchange NASA data and information. Currently, ECHO has metadata descriptors for over 55 million individual data granules and 13 million browse images. The majority of ECHO's holdings come directly from data held in the NASA DAACs. The science disciplines and domains represented in ECHO are diverse and include metadata for all of NASA's Science Focus Area data. As middleware for a service-oriented enterprise, ECHO offers access to its capabilities through a set of publicly available Application Program Interfaces (APIs). More information about ECHO is available at http://eos.nasa.gov.echo. The presentation will discuss the status of the ECHO Partners, holdings, and activities, including the transition from the EOS Data Gateway to the Warehouse Inventory Search Tool (WIST)

Weinstein, Beth↗

Simple, Script-Based Science Processing Archive

The Simple, Scalable, Script-based Science Processing (S4P) Archive (S4PA) is a disk-based archival system for remote sensing data. It is based on the data-driven framework of S4P and is used for data transfer, data preprocessing, metadata generation, data archive, and data distribution. New data are automatically detected by the system. S4P provides services such as data access control, data subscription, metadata publication, data replication, and data recovery. It comprises scripts that control the data flow. The system detects the availability of data on an FTP (file transfer protocol) server, initiates data transfer, preprocesses data if necessary, and archives it on readily available disk drives with FTP and HTTP (Hypertext Transfer Protocol) access, allowing instantaneous data access. There are options for plug-ins for data preprocessing before storage. Publication of metadata to external applications such as the Earth Observing System Clearinghouse (ECHO) is also supported. S4PA includes a graphical user interface for monitoring the system operation and a tool for deploying the system. To ensure reliability, S4P continuously checks stored data for integrity, Further reliability is provided by tape backups of disks made once a disk partition is full and closed. The system is designed for low maintenance, requiring minimal operator oversight.

Lynnes, Christopher↗

MoonDB: Restoration and Synthesis of Lunar Petrological and Geochemical Data

About 2,200 samples were collected from the Moon during the Apollo missions, forming a unique and irreplaceable legacy of the Apollo program. These samples, obtained at tremendous cost and great risk, are the only samples that have ever been returned by astronauts from the surface of another planetary body. These lunar samples have been curated at NASA Johnson Space Center and made available to the global research community. Over more than 45 years, a vast body of petrological, geochemical, and geochronological studies of these samples have been amassed, which helped to expand our understanding of the history and evolution of the Moon, the Earth itself, and the history of our entire solar system. Unfortunately, data from these studies are dispersed in the literature, often only available in analog format in older publications, and/or lacking sample metadata and analytical metadata (e.g., information about analytical procedure and data quality), which greatly limits their usage for new scientific endeavors. Even worse is that much lunar data have never been published, simply because no forum existed at the time (e.g., electronic supplements). Thousands of valuable analyses remain inaccessible, often preserved only in personal records, and are in danger of being lost forever, when investigators retire or pass away. Making these data and metadata publicly accessible in a digital format would dramatically help guide current and future research and eliminate duplicated analyses of precious lunar samples.

Lehnert, Kerstin A.↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗

Extending the Reach of IGSN Beyond Earth: Implementing IGSN Registration to Link Nasa's Apollo Lunar Samples and Their Data

The rock and soil samples returned from the Apollo missions from 1969-72 have supported 46 years of research leading to advances in our understanding of the formation and evolution of the inner Solar System. NASA has been engaged in several initiatives that aim to restore, digitize, and make available to the public existing published and unpublished research data for the Apollo samples. One of these initiatives is a collaboration with IEDA (Interdisciplinary Earth Data Alliance) to develop MoonDB, a lunar geochemical database modeled after PetDB (Petrological Database of the Ocean Floor). In support of this initiative, NASA has adopted the use of IGSN (International Geo Sample Number) to generate persistent, unique identifiers for lunar samples that scientists can use when publishing research data. To facilitate the IGSN registration of the original 2,200 samples and over 120,000 subdivided samples, NASA has developed an application that retrieves sample metadata from the Lunar Curation Database and uses the SESAR API to automate the generation of IGSNs and registration of samples into SESAR (System for Earth Sample Registration). This presentation will describe the work done by NASA to map existing sample metadata to the IGSN metadata and integrate the IGSN registration process into the sample curation workflow, the lessons learned from this effort, and how this work can be extended in the future to help deal with the registration of large numbers of samples.

Todd, Nancy S.↗

Heuristics for Relevancy Ranking of Earth Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗

Relevancy Ranking of Satellite Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗