Search NASASearch

SEARCH · Search NASA

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors

Development of an Improved Spatial Metadata Simplification Algorithm

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata

Use of Spatial Metadata Simplification for TEMPO

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata

Logic programming and metadata specifications

Artificial intelligence (AI) ideas and techniques are critical to the development of intelligent information systems that will be used to collect, manipulate, and retrieve the vast amounts of space data produced by 'Missions to Planet Earth.' Natural language processing, inference, and expert systems are at the core of this space application of AI. This paper presents logic programming as an AI tool that can support inference (the ability to draw conclusions from a set of complicated and interrelated facts). It reports on the use of logic programming in the study of metadata specifications for a small problem domain of airborne sensors, and the dataset characteristics and pointers that are needed for data access.

Lopez, Antonio M., Jr.

The VIS-AD data model: Integrating metadata and polymorphic display with a scientific programming language

The VIS-AD data model integrates metadata about the precision of values, including missing data indicators and the way that arrays sample continuous functions, with the data objects of a scientific programming language. The data objects of this data model form a lattice, ordered by the precision with which they approximate mathematical objects. We define a similar lattice of displays and study visualization processes as functions from data lattices to display lattices. Such functions can be applied to visualize data objects of all data types and are thus polymorphic.

Hibbard, William L.

Textural-Contextual Labeling and Metadata Generation for Remote Sensing Applications

Despite the extensive research and the advent of several new information technologies in the last three decades, machine labeling of ground categories using remotely sensed data has not become a routine process. Considerable amount of human intervention is needed to achieve a level of acceptable labeling accuracy. A number of fundamental reasons may explain why machine labeling has not become automatic. In addition, there may be shortcomings in the methodology for labeling ground categories. The spatial information of a pixel, whether textural or contextual, relates a pixel to its surroundings. This information should be utilized to improve the performance of machine labeling of ground categories. Landsat-4 Thematic Mapper (TM) data taken in July 1982 over an area in the vicinity of Washington, D.C. are used in this study. On-line texture extraction by neural networks may not be the most efficient way to incorporate textural information into the labeling process. Texture features are pre-computed from cooccurrence matrices and then combined with a pixel's spectral and contextual information as the input to a neural network. The improvement in labeling accuracy with spatial information included is significant. The prospect of automatic generation of metadata consisting of ground categories, textural and contextual information is discussed.

Kiang, Richard K.

Data Archival and Retrieval Enhancement (DARE) Metadata Modeling and Its User Interface

The Defense Nuclear Agency (DNA) has acquired terabytes of valuable data which need to be archived and effectively distributed to the entire nuclear weapons effects community and others...This paper describes the DARE (Data Archival and Retrieval Enhancement) metadata model and explains how it is used as a source for generating HyperText Markup Language (HTML)or Standard Generalized Markup Language (SGML) documents for access through web browsers such as Netscape.

The Defense Nuclear Agency DNA DARE Data Archival

Availability of Previously Unprocessed ALSEP Raw Instrument Data, Derivative Data, and Metadata Products

In year 2010, 440 original data archival tapes for the Apollo Lunar Science Experiment Package (ALSEP) experiments were found at the Washington National Records Center. These tapes hold raw instrument data received from the Moon for all the ALSEP instruments for the period of April through June 1975. We have recently completed extraction of binary files from these tapes, and we have delivered them to the NASA Space Science Data Cordinated Archive (NSSDCA). We are currently processing the raw data into higher order data products in file formats more readily usable by contemporary researchers. These data products will fill a number of gaps in the current ALSEP data collection at NSSDCA. In addition, we have estabilished a digital, searcheable archive of ALSEP document and metadata as part of the web portal of the Lunar and Planetary Institute. It currently holds approx. 700 documents totaling approx. 40,000 pages

ALSEP

Storage of Physical Sample Metadata in the Astrobiology Habitable Environments Database (AHED)

The National Aeronautics and Space Administration has begun an effort to store, curate, and publish information about physical samples collected and analyzed in conjunction with NASA-funded astrobiology research. Astrobiology is a multidisciplinary area of scientific research being conducted by collaborating teams of biologists, chemists, geologists, atmospheric scientists, oceanographers, astrophysicists, astronomers, and other specialists. Astrobiology studies the origin, evolution, and distribution of life in the Universe. NASA uses the results of astrobiology research to focus its future missions on targets of opportunity for the discovery of life off Earth. Astrobiology researchers conduct both field-based and laboratory-based research, during which physical samples are collected, processed, and catalogued. The cataloguing practices employed by different teams of astrobiologists vary widely, and there are no specific standards available to guide the collection and recording of astrobiology sample data. The disparity in data collection approaches and the lack of a centralized sample repository makes it difficult for astrobiology teams to share data and benefit from resultant synergies.To facilitate data sharing within the astrobiology community, NASA is developing a prototype database the Astrobiology Habitable Environments Database (AHED) and an associated set of data collection templates. The database will store information about samples, along with associated measurements and analyses, including information about biological cultures enriched or isolated from samples, and the results of analyses performed on the samples (e.g., via spectrography, microscopy, etc.). In addition, the system will store contextual information about field sites where samples were collected, the instruments or equipment used for analysis, and people and institutions involved in their collection. AHED is being implemented on top of Open Data Repository's Data Publisher [1], an open source software platform for the publication of scientific datasets. The data collection templates under development represent an initial attempt to propose a set of metadata for capture and storage within AHED. The design of these templates is being conducted by a consolidated group of astrobiologists from active research teams at NASA Ames Research Center, assisted by data science and software engineering specialists. These initial templates must be vetted with the broader astrobiology community through a defined process to ensure that they meet community needs. Each template captures a different type of data collection record. For each template, we are developing a list of fields to be captured, including a set of required entry fields, a set of recommended but optional fields, and a set of discretionary fields. A datatype selected from a variety of text and numeric types is specified for each field. Included is a 'choice' type that restricts user input to an enumerated list of values. Many of the fields and field values capture information of particular interest to the astrobiology community, and are intended to facilitate search and retrieval of relevant data across multiple datasets.

Keller, Rich

GeneLab Metadata & Processed Data

An overview of the organization and structure of the metadata and data in the GeneLab Data Repository. This presentaiton provides examples of how the data is presented and what data can be download from the GLDS Repository.

Gebre, Sam

Availability of previously lost data and metadata from the Apollo Lunar Surface Experiments Package (ALSEP)

Fourteen types of geophysical instruments deployed at the Apollo 12, 14, 15, 16, and 17 sites by the astronauts for long-term observation were collectively called the Apollo Lunar Surface Experiments Package (ALSEP). These instruments were active from the times of their deployment (November 1969–December 1972) to September 1977. At the conclusion of the experiments, the raw instrument data received from the Moon prior to March 1976 were left unarchived. Portions of the data processed by the principal investigators (PIs) of these experiments had been archived at the NASA Space Science Data Coordinated Archive (NSSDCA) in various formats. The unarchived data, residing then on open-reel magnetic tapes, became lost in the decades since, along with much of the metadata (the supporting documents for these data). We have recently recovered 440 of the previously lost tapes, containing raw ALSEP instrument data from April through June of 1975. Here we describe the data extracted from these tapes and summarize the data products generated for archiving at the NASA Planetary Data System (PDS) and NSSDCA, along with their historical narrative. In addition, we have reformatted many of the datasets delivered to NSSDCA by the PIs in the 1970s for archiving at the PDS. Finally, we have compiled an online searchable repository of ALSEP-related documents by optically scanning tens of thousands of pages of them kept at the Lunar and Planetary Institute in Texas.

S. Nagihara

Increasing Accessibility of the Runs-on-Request Metadata, Data, and Services at the Community Coordinated Modeling Center

Space weather models are essential to our ability to understand and predict space weather events. For over 20 years, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) has been providing transformative tools and platforms for hosting space weather models and associated services, free and open to anyone interested in studying space weather. Runs-on-Request system (ROR) is one of the popular services at CCMC that permits researchers and other end-users to exercise cutting-edge hosted heliophysics and space weather models using a simple web interface, as well as collaborate on an extensive and continuously growing archive of over 28,000 model run results. Similar to other projects at CCMC, ROR has grown as a community project that strives to be open and transparent to its users. In this poster, we discuss some of our recent efforts to further expose ROR data, metadata, and services to the end users through both custom and community-developed access protocols. We also discuss how in-house science support provided by the CCMC team plays a paramount role in making ROR data and services truly accessible by the community.

Maksym Petrenko

Metadata Entry Optimization for NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Sample Repository

Metadata Entry Optimization For NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Biospecimen