Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Systems; Data Services; Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Newly Developing Community-Oriented Data System from NASA GES DISC

Data services are essential to facilitate data access and to aid efficiency of conducting research and application activities. With emerging technologies such as cloud computing and AI/ML (Artificial Intelligence/Machine Learning) leading the pace of the data world, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), home to the permanent archive for multidisciplinary Earth Observation (EO) geospatial data to study atmospheric composition, weather and climate variability, and water and energy cycles is no exception.Interfacing directly with users as part of data center work, we understand the challenges for the required time and effort to discover, visualize, and analyze large varieties and quantities of Earth Observation information for research, monitoring, and decision-making, largely due to the existing data and information systems aim to support experienced users, but has been proved difficult for non-earth scientists and new users that are unfamiliar with the variety of formats and structures in which data, metadata, and information are stored, as well as the required methods to use them. To address these challenges, I will update our latest activities with regard to water-and energy-related products and community-oriented and user-friendly services at the GES DISC, including our plans for the emerging technologies.

Jennifer Wei↗

Recovering Nimbus Era Observations at the NASA GES DISC

Between 1964 and 1978, NASA launched a series of seven Nimbus meteorological satellites which provided Earth observations for 30 years. These satellites, carrying a total of 33 instruments to observe the Earth at visible, infrared, ultraviolet, and microwave wavelengths, revolutionized weather forecasting, provided early observations of ocean color and atmospheric ozone, and prototyped location-based search and rescue capabilities. The Nimbus series paved the way for a number of currently operational systems such as the EOS (Earth Observation System) Terra, Aqua, and Aura platforms. The original data archive includes both magnetic tapes and film media. These media are well past their expected end of life, placing at risk valuable data that are critical to extending the history of Earth observations back in time. GES DISC (Goddard Earth Sciences Data and Information Services Center) has been incorporating these data into a modern online archive by recovering the digital data files from the tapes, and scanning images of the data from film strips. The digital data products were written on obsolete hardware systems in outdated file formats, and in the absence of metadata standards at that time, were often written in proprietary file structures. Through a tedious and laborious process, oft-corrupted data are recovered, and incomplete metadata and documentation are reconstructed.

data recovery↗

Enabling Cloud Services and Enhanced Data Discovery With Earthdata-Varinfo

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. Each of these collections can contain hundreds of variables that describe each measurement within the dataset, therefore an automated method for generating UMM-Var records is necessary. The Unified Metadata Model for Variables (UMM-Var) provides a framework for variable metadata records in NASA’s Common Metadata Repository (CMR). The Python tool, earthdata-varinfo, was developed to solve this problem of automating the curation of UMM-Var records. Given either a collection DMR file or a netCDF-4 file, earthdata-varinfo can scrape variable metadata and return a CMR compliant UMM-Var record. Earthdata-varinfo can generate thousands of UMM-Var records in a matter of seconds, thus enabling subsetting capabilities and enhancing data discovery.

Eni Awowale↗

Proto-Examples of Data Access and Visualization Components of a Potential Cloud-Based GEOSS-AI System

Once a research or application problem has been identified, one logical next step is to search for available relevant data products. Thus, an early component of a potential GEOSS-AI system, in the continuum between observations and end point research, applications, and decision making, would be one that enables transparent data discovery and access by users. Such a component might be effected via the systems data agents. Presumably, some kind of data cataloging has already been implemented, e.g., in the GEOSS Common Infrastructure (GCI). Both the agents and cataloging could also leverage existing resources external to the system. The system would have some means to accept and integrate user-contributed agents. The need or desirability for some data format internal to the system should be evaluated. Another early component would be one that facilitates browsing visualization of the data, as well as some basic analyses.Three ongoing projects at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) provide possible proto-examples of potential data access and visualization components of a cloud-based GEOSS-AI system. 1. Reorganizing data archived as time-step arrays to point-time series (data rods), as well as leveraging the NASA Simple Subset Wizard (SSW), to significantly increase the number of data products available, at multiple NASA data centers, for production as on-the-fly (virtual) data rods. SSWs data discovery is based on OpenSearch. Both pre-generated and virtual data rods are accessible via Web services. 2. Developing Web Feature Services to publish the metadata, and expose the locations, of pre-generated and virtual data rods in the GEOSS Portal and enable direct access of the data via Web services. SSW is also leveraged to increase the availability of both NASA and non-NASA data.3.Federating NASA Giovanni (Geospatial Interactive Online Visualization and Analysis Interface), for multi-sensor data exploration, that would allow each cooperating data center, currently the NASA Distributed Active Archive Centers (DAACs), to configure its own Giovanni deployment, while also allowing all the deployments to incorporate each others data. A federated Giovanni comprises Giovanni Virtual Machines, which can be run on local servers or in the cloud.

access↗

Finding Atmospheric Composition (AC) Metadata

The Atmospheric Composition Portal (ACP) is an aggregator and curator of information related to remotely sensed atmospheric composition data and analysis. It uses existing tools and technologies and, where needed, enhances those capabilities to provide interoperable access, tools, and contextual guidance for scientists and value-adding organizations using remotely sensed atmospheric composition data. The initial focus is on Essential Climate Variables identified by the Global Climate Observing System CH4, CO, CO2, NO2, O3, SO2 and aerosols. This poster addresses our efforts in building the ACP Data Table, an interface to help discover and understand remotely sensed data that are related to atmospheric composition science and applications. We harvested GCMD, CWIC, GEOSS metadata catalogs using machine to machine technologies - OpenSearch, Web Services. We also manually investigated the plethora of CEOS data providers portals and other catalogs where that data might be aggregated. This poster is our experience of the excellence, variety, and challenges we encountered.Conclusions:1.The significant benefits that the major catalogs provide are their machine to machine tools like OpenSearch and Web Services rather than any GUI usability improvements due to the large amount of data in their catalog.2.There is a trend at the large catalogs towards simulating small data provider portals through advanced services. 3.Populating metadata catalogs using ISO19115 is too complex for users to do in a consistent way, difficult to parse visually or with XML libraries, and too complex for Java XML binders like CASTOR.4.The ability to search for Ids first and then for data (GCMD and ECHO) is better for machine to machine operations rather than the timeouts experienced when returning the entire metadata entry at once. 5.Metadata harvest and export activities between the major catalogs has led to a significant amount of duplication. (This is currently being addressed) 6.Most (if not all) Earth science atmospheric composition data providers store a reference to their data at GCMD.

metadata search↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗

Digital Archive Issues from the Perspective of an Earth Science Data Producer

Contents include the following: Introduction. A Producer Perspective on Earth Science Data. Data Producers as Members of a Scientific Community. Some Unique Characteristics of Scientific Data. Spatial and Temporal Sampling for Earth (or Space) Science Data. The Influence of the Data Production System Architecture. The Spatial and Temporal Structures Underlying Earth Science Data. Earth Science Data File (or Relation) Schemas. Data Producer Configuration Management Complexities. The Topology of Earth Science Data Inventories. Some Thoughts on the User Perspective. Science Data User Communities. Spatial and Temporal Structure Needs of Different Users. User Spatial Objects. Data Search Services. Inventory Search. Parameter (Keyword) Search. Metadata Searches. Documentation Search. Secondary Index Search. Print Technology and Hypertext. Inter-Data Collection Configuration Management Issues. An Archive View. Producer Data Ingest and Production. User Data Searching and Distribution. Subsetting and Supersetting. Semantic Requirements for Data Interchange. Tentative Conclusions. An Object Oriented View of Archive Information Evolution. Scientific Data Archival Issues. A Perspective on the Future of Digital Archives for Scientific Data. References Index for this paper.

Barkstrom, Bruce R.↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database↗

Progress and Plans in Support of the Polar Community

Feedback provided by the Antarctic community has proven instrumental in positively influencing the direction of the GCMD's development. For example, in response to requests for a stand alone metadata authoring tool, a new shareable software package called docBUILDER solo will be released to the public in March 2006. This tool permits researchers to document their data during experiments and observational periods in the field. The international polar community has also played a key role in encouraging support for the foreign language character set in the metadata display and tools (10% of the records in the AMD hold foreign characters). In the upcoming release, the full ISO character set, which also includes mathematical symbols, will be supported. Additional upgrades include the ability for users to search for data sets based on pre-selected temporal and spatial resolution ranges. Data providers are strongly encouraged to populate the resolution fields for their data sets, although these fields are not currently required. In prior versions, browser incompatibilities often resulted in unreliable performance for users attempting to initiate a spatial search using a map based on Java applet technology. The GCMD will offer an integrated Google map and date search, replacing the applet technology and enhancing the geospatial and temporal searches. It is estimated that 30% of the records in the AMD have direct access to data. A growing number of these records can be accessed through data service links. Related data services are therefore becoming valuable assets in facilitating the use and visualization of data. Users will gain the ability to refine services using the same options as those available for data set searches. Data providers are encouraged to describe available data-related services through the directory. Future plans include offering web services through a SOAP interface and extending semantic queries for the polar regions through the use of ontologies. The Open Archives Initiative's (OAI) Protocol for Metadata Harvesting (PMH) has been successfully tested with several organizations and appears to be a prime candidate for sharing metadata within the community. The GCMD anticipates contributing to the design of the data management system for the International Polar Year and to the ongoing efforts in the years to come. Further enhancements will be discussed at the meeting.

Olsen, Lola M.↗

Emerging Network Storage Management Standards for Intelligent Data Storage Subsystems

This paper discusses the need for intelligent storage devices and subsystems that can provide data integrity metadata, the content of the existing data integrity standard for optical disks and techniques and metadata to verify stored data on optical tapes developed by the Association for Information and Image Management (AIIM) Optical Tape Committee.

OPTICAL DATA STORAGE MATERIALS↗

New Metadata Capabilities within NASA's Common Metadata Repository (CMR)

This talk will convey the new capabilities and features of the UMM-Variables and the UMM-Services metadata model within the Common Metadata Repository (CMR) and how community engagement through the ESDIS Standards Office (ESO) review shaped the new versions of the models. The Unified Metadata Model (UMM) provides a common metadata model to unify legacy systems (i.e. GCMD (Global Change Master Directory), ECHO (Earth Observing System (EOS) Clearinghouse)) with new systems (i.e. CMR). The rationale and migration process of the Service Entry Resource Format (SERF) to the UMM-S will also be conveyed. The talk will conclude with discussing issues and lessons learned from the review process and how future reviews will be conducted to ensure a more targeted, meaningful review with a faster turn-around time for triage and implementation.

Earth Observing System↗

EOSDIS STAC Briefing

The Spatio-Temporal Asset Catalog (STAC) specification provides a common language to describe a range of geospatial information, so it can more easily be indexed and discovered. A 'spatiotemporal asset' is any file that represents information about the earth captured in a certain space and time. STAC provides a standards-based, web and cloud friendly cataloging specification that has seen a large amount of adoption in the cloud geospatial information system space. This presentation will demonstrate how NASA EOSDIS leverages STAC technology to provide value-added features to both our data discovery and transformation services. Finally, we will propose a way to improve our federated discovery capabilities using STAC.

NASA EOSDIS STAC Cloud catalog metadata↗

Towards the Development of a Unified Distributed Date System for L1 Spacecraft

The purpose of this grant, 'Towards the Development of a Unified Distributed Data System for L1 Spacecraft', is to take the initial steps towards the development of a data distribution mechanism for making in-situ measurements more easily accessible to the scientific community. Our obligations as subcontractors to this grant are to add our Faraday Cup plasma data to this initial study and to contribute to the design of a general data distribution system. The year 1 objectives of the overall project as stated in the GSFC proposal are: 1) Both the rsync and Perl based data exchange tools will be fully developed and tested in our mixed, Unix, VMS, Windows and Mac OS X data service environment. Based on the performance comparisons, one will be selected and fully deployed. Continuous data exchange between all L1 solar wind monitors initiated. 2) Data version metadata will be agreed upon, fully documented, and deployed on our data sites. 3) The first version of the data description rules, encoded in a XML Schema, will be finalized. 4) Preliminary set of library routines will be collected, documentation standards and formats agreed on, and desirable routines that have not been implemented identified and assigned. 5) ViSBARD test site implemented to independently validate data mirroring procedures. The specific MIT tasks over the duration of this project are the following: a) implement mirroring service for WIND plasma data b) participate in XML Schema development c) contribute toward routine library.

Lazarus, Alan J.↗

WGISS CWIC Report: CWIC Evolution

The purpose of the Committee on Earth Observing Satellites (CEOS) Working Group on Information Systems and Services (WGISS) Integrated Catalog (CWIC) is to provide a consistent search interface to help users find and access satellite data made available by CWIC data partners through the use of the WGISS-supported standards. This presentation will cover the current status of CWIC, CWIC Data partners, CWIC Metrics, and the CWIC evolution/transition activities.

metadata↗

COMPASS-FME Synoptic Sites Level 1 Sensor Data v2-1

This is the version 2-1 Level 1 (L1) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems. L1 data are close to raw, but are units-transformed and have out-of-instrument-bounds, out-of-service, and outlier flags added. Duplicates and missing data are removed but otherwise these data are not filtered, and have not been subject to any additional algorithmic or human QA/QC. Any scientific analyses of L1 data should be performed with care. **This dataset will be updated quarterly with new data for the duration of the project** This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding up to 12 CSV (comma separated value) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are normally logged every 15 minutes. Please see v2-0 Synoptic L1 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. This dataset was updated 2026-03-12: (i) data now go through 2025-12-31 (previous end was 2025-06-30) and (ii) dataset and file names updated to “…v2-1” (previously was “v2-0”).

54 ENVIRONMENTAL SCIENCES↗

Preserving NASA Historic and Current Mission Data and Adding Value to These for Future Researchers

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has been actively involved in many aspects of ensuring the long-term preservation of NASA earth science data and knowledge. This involves both the recovery and preservation of early NASA meteorological and other earth observation data, as well as preserving the more recent Earth Observation System (EOS) mission data sets which continue or have reached their end of lifetime. The GES DISC adds value to these preserved data by adding metadata and making the data available online to future researchers. The early NASA meteorological and earth observation data sets from the 1960s and 70s were originally archived on magnetic tapes, and visualizations of these data were preserved on 70-mm film. As these media have aged, their contents have been at risk of permanent loss. NASA has given the task of preserving these early data sets to the GES DISC and making these data sets easily available to the public. The data from these early missions are potentially useful to climate researchers as these are some of the only global measurements made at their time. These old data on magnetic tapes and film strips do not contain easily readable metadata, and so to add value the GES DISC has added digital metadata to them so that the data are searchable and findable. The GES DISC is also involved in preserving the data and knowledge from the EOS era missions. The GES DISC follows the guidelines developed for the preservation of data as specified in the NASA EOS Data and Information System (EOSDIS) Earth Science Data Preservation Content Specification (423-SPEC-001) document. To date, the GES DISC has consulted with the data science teams from the following missions: UARS, Earth Probe TOMS, Aura HIRDLS, and SORCE, in order to properly preserve their data and accompanying documentation. The GES DISC is also currently working with the EOS science teams from TRMM, AIRS, MLS, OMI and additional missions to ensure that the relevant documents and data sets are properly archived for future researchers. A standardized procedure for mission data preservation following 423-SPEC-001 makes preservation among the many NASA EOSDIS data centers uniform, so that these could be transitioned easily to a common EOSDIS preservation repository. This presentation will give an overview of the preservation and recovery of the old NASA historical data sets archived at the GES DISC, as well as the data and documentation preservation efforts of the EOS era missions.

James Johnson↗

COMPASS-FME Synoptic Sites Level 2 Sensor Data v2-1

This is the version 2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. Please see v2-1 L2 Sensor Package QStart.pdf for detailed information on data package structure, temporal coverage, and versioning.

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗