Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

A Newly Developing Community-Oriented Data System from NASA GES DISC

Data services are essential to facilitate data access and to aid efficiency of conducting research and application activities. With emerging technologies such as cloud computing and AI/ML (Artificial Intelligence/Machine Learning) leading the pace of the data world, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), home to the permanent archive for multidisciplinary Earth Observation (EO) geospatial data to study atmospheric composition, weather and climate variability, and water and energy cycles is no exception.Interfacing directly with users as part of data center work, we understand the challenges for the required time and effort to discover, visualize, and analyze large varieties and quantities of Earth Observation information for research, monitoring, and decision-making, largely due to the existing data and information systems aim to support experienced users, but has been proved difficult for non-earth scientists and new users that are unfamiliar with the variety of formats and structures in which data, metadata, and information are stored, as well as the required methods to use them. To address these challenges, I will update our latest activities with regard to water-and energy-related products and community-oriented and user-friendly services at the GES DISC, including our plans for the emerging technologies.

Jennifer Wei↗

Limiting Data Friction by Reducing Data Download Using Spatiotemporally Aligned Data Organization Through STARE

Current data processing practice limits the volume and variety of relevant geoscience data that can practically be applied to important problems. File archives in centralized data centers are the principal means by which Earth Science data are accessed. This approach, however, requires laborious search, retrieval, and eventual customization/adaptation for the data to be used. Such fractionation makes it even more difficult to share outcomes, i.e. research artifacts and data products, hampering reusability and repeatability, since end users generally have their own research agenda and preferences as well as scarce resources. Thus, while finding and downloading data files from central data centers are already costly for end users working in their own field, using data products from other disciplines rapidly becomes prohibitive. This curtails scientific productivity, limits avenues of study, and endangers quality and reproducibility. The Spatio-Temporal Adaptive Resolution Encoding (STARE) is a unifying scheme that facilitates the indexing, access, and fusion of diverse Earth Science data. STARE implements an innovative encoding of geo-spatiotemporal information, originally developed for aligning datasets with diverse spatiotemporal characteristics in an array database. The spatial component of STARE recursively quadfurcates a root polyhedron, producing a hierarchical scheme for addressing geographic locations and regions. The temporal component of STARE uses conventional date-time units as an indexing hierarchy. The additional encoding of spatial and temporal resolution information in STARE enables comparisons and conditional selections across diverse datasets. Moreover, spatiotemporal set-operations, e.g. union and intersection, are mapped to efficient integer operations with STARE. Applied to existing data models (point, grid, spacecraft swath) and corresponding granules, STARE indexes provide a streamlined description usable as geo-spatiotemporal metadata. When coupled with large scale, distributed hardware and software, STARE-based data access reduces pre-analysis data preparation costs by offering a convenient means to align different datasets spatiotemporally without specialized effort in parallel computing or distributed data management.

Kuo, Kwo-Sen↗

Improving Earth Science Dataset Search with Publication

The NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) archives a large number of Earth observational datasets. Thousands of the publications are created each year based on these datasets. The content of these publications can be used for discovery of the datasets based on the characteristics of applicational research. We leverage the content of these publications to retrieve the information about phenomena and domains where measurements from the datasets were utilized through linking these publications and dataset in Knowledge Graph. We retrieve phenomena and domain information using SWEET ontology and produce the set of keywords that are linked to the datasets. Further, we evaluate this link strength according to the frequency of dataset usage in the papers mentioning these keywords. We demonstrate how this linkage can improve dataset search by comparing the search results obtained from Common Metadata Repository (CMR) search and the publications based data.

Kristina Stoyanova↗

Working at NASA as a (Molecular) Biologist...

Trained as a biologist, I landed my first job at NASA as a mission scientist on the Rodent Research project. As a mission scientist, I led science activities associated with development, flight, and post-flight analysis of rodent research missions on the ISS in consultation with project engineers, operational specialists, mission PI(s), and mission Implementation Partner(s). I also assessed the feasibility of scientific experiments as a technical and scientific expert in rodent research on the ISS. Eager to learn more about the molecular changes that occur as a result of spaceflight exposure, I transitioned to the NASA GeneLab project where I eventually became the data processing lead. NASA’s GeneLab helps scientists understand how the fundamental building blocks of life – DNA, RNA, proteins, and metabolites – change from exposure to the space environment including microgravity and cosmic radiation exposure. GeneLab does so by providing fully coordinated epigenomics, genomics, transcriptomics, proteomics, and metabolomics data (collectively known as omics data) alongside essential metadata describing each spaceflight and space-relevant experiment. In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab has started processing and analyzing these datasets to generate differential gene expression data and identify biological and physiological pathways that are dysregulated as a result of spaceflight. The user interface was designed to be accessible to a broad variety of users, including high school and college students who can use it to learn about omics data analysis and space biology. During the CA Space Grant NASA Panel event, I will provide an overview of my work at NASA on both the Rodent Research and GeneLab projects and will conclude by providing resources for opportunities to work with GeneLab and NASA at large.

Amanda M Saravia-Butler↗

My Journey to NASA and role as a (Molecular) Biologist…

My journey to NASA started in undergrad, where I performed microbiology research and became hooked in the STEM fields. From there I earned a Ph.D. in Biochemistry and Molecular Biology from Mayo Graduate school where I performed Pancreatic Cancer research and then went on to perform postdoctoral research in Developmental Biology at the University of Miami. Trained as a biologist, I landed my first job at NASA as a mission scientist on the Rodent Research project. As a mission scientist, I led science activities associated with development, flight, and post-flight analysis of rodent research missions on the ISS in consultation with project engineers, operational specialists, mission PI(s), and mission Implementation Partner(s). I also assessed the feasibility of scientific experiments as a technical and scientific expert in rodent research on the ISS. Eager to learn more about the molecular changes that occur as a result of spaceflight exposure, I transitioned to the NASA GeneLab project where I eventually became the data processing lead. NASA’s GeneLab helps scientists understand how the fundamental building blocks of life – DNA, RNA, proteins, and metabolites – change from exposure to the space environment including microgravity and cosmic radiation exposure. GeneLab does so by providing fully coordinated epigenomics, genomics, transcriptomics, proteomics, and metabolomics data (collectively known as omics data) alongside essential metadata describing each spaceflight and space-relevant experiment. To maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab has started processing and analyzing these datasets to generate differential gene expression data and identify biological and physiological pathways that are dysregulated as a result of spaceflight. The user interface was designed to be accessible to a broad variety of users, including high school and college students who can use it to learn about omics data analysis and space biology. During the CA Space Grant Women in STEM Panel event, I will provide an overview of my journey to NASA including my work at NASA on both the Rodent Research and GeneLab projects. I will conclude by providing resources for opportunities to work with GeneLab and NASA at large followed by links to programs designed specifically to engage women in STEM fields.

Amanda M Saravia-Butler↗

My Journey to NASA and Role as a (Molecular) Biologist…

My journey to NASA started in undergrad, where I performed microbiology research and became hooked in the STEM fields. From there I earned a Ph.D. in Biochemistry and Molecular Biology from Mayo Graduate school where I performed Pancreatic Cancer research and then went on to perform postdoctoral research in Developmental Biology at the University of Miami. Trained as a biologist, I landed my first job at NASA as a mission scientist on the Rodent Research project. As a mission scientist, I led science activities associated with development, flight, and post-flight analysis of rodent research missions on the ISS in consultation with project engineers, operational specialists, mission PI(s), and mission Implementation Partner(s). I also assessed the feasibility of scientific experiments as a technical and scientific expert in rodent research on the ISS. Eager to learn more about the molecular changes that occur as a result of spaceflight exposure, I transitioned to the NASA GeneLab project where I eventually became the data processing lead. NASA’s GeneLab helps scientists understand how the fundamental building blocks of life – DNA, RNA, proteins, and metabolites – change from exposure to the space environment including microgravity and cosmic radiation exposure. GeneLab does so by providing fully coordinated epigenomics, genomics, transcriptomics, proteomics, and metabolomics data (collectively known as omics data) alongside essential metadata describing each spaceflight and space-relevant experiment. To maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab has started processing and analyzing these datasets to generate differential gene expression data and identify biological and physiological pathways that are dysregulated as a result of spaceflight. The user interface was designed to be accessible to a broad variety of users, including high school and college students who can use it to learn about omics data analysis and space biology. During the NASA Ames ECN at Lincoln HS Virtual Panel event, I will provide an overview of my journey to NASA including my work at NASA on both the Rodent Research and GeneLab projects. I will conclude by providing resources for opportunities to work with GeneLab and NASA at large followed by links to programs designed specifically to engage women in STEM fields.

Amanda M Saravia-Butler↗

HAPI: An API Standard for Accessing Heliophysics Time Series Data

Heliophysics data analysis often involves combining diverse science measurements, many of them captured as time series. Although there are now only a few commonly used data file formats, the diversity in mechanisms for automated access to and aggregation of such data holdings can make analysis that requires intercomparison of data from multiple data providers difficult. The Heliophysics Application Programmer's Interface (HAPI) is a recently developed standard for accessing distributed time series data to increase interoperability. The HAPI specification is based on the common elements of existing data services, and it standardizes the two main parts of a data service: the request interface and the response data structures. The interface is based on the REpresentational State Transfer (REST) or RESTful architecture style, and the HAPI specification defines five required REST endpoints. Data are returned via a streaming format that hides file boundaries; the metadata is detailed enough for the content to be scientifically useful, e.g., plotted with appropriate axes layout, units, and labels. Multiple mature HAPI-related open-source projects offer server-side implementation tools and client-side libraries for reading HAPI data in multiple languages (IDL, Java, MATLAB, and Python). Multiple data providers in the US and Europe have added HAPI access alongside their existing interfaces. Based on this experience, data can be served via HAPI with little or no information loss compared to similar existing web interfaces. Finally, HAPI has been recommended as a COSPAR standard for time series data delivery.

Robert S. Weigel↗

Planetary Image Geometry Library

The Planetary Image Geometry (PIG) library is a multi-mission library used for projecting images (EDRs, or Experiment Data Records) and managing their geometry for in-situ missions. A collection of models describes cameras and their articulation, allowing application programs such as mosaickers, terrain generators, and pointing correction tools to be written in a multi-mission manner, without any knowledge of parameters specific to the supported missions. Camera model objects allow transformation of image coordinates to and from view vectors in XYZ space. Pointing models, specific to each mission, describe how to orient the camera models based on telemetry or other information. Surface models describe the surface in general terms. Coordinate system objects manage the various coordinate systems involved in most missions. File objects manage access to metadata (labels, including telemetry information) in the input EDRs and RDRs (Reduced Data Records). Label models manage metadata information in output files. Site objects keep track of different locations where the spacecraft might be at a given time. Radiometry models allow correction of radiometry for an image. Mission objects contain basic mission parameters. Pointing adjustment ("nav") files allow pointing to be corrected. The object-oriented structure (C++) makes it easy to subclass just the pieces of the library that are truly mission-specific. Typically, this involves just the pointing model and coordinate systems, and parts of the file model. Once the library was developed (initially for Mars Polar Lander, MPL), adding new missions ranged from two days to a few months, resulting in significant cost savings as compared to rewriting all the application programs for each mission. Currently supported missions include Mars Pathfinder (MPF), MPL, Mars Exploration Rover (MER), Phoenix, and Mars Science Lab (MSL). Applications based on this library create the majority of operational image RDRs for those missions. A Java wrapper around the library allows parts of it to be used from Java code (via a native JNI interface). Future conversions of all or part of the library to Java are contemplated.

Deen, Robert C.↗

Quantitative Comparison of Proprietary and Open-Source Georeferencing Tools for Use with Astronaut Photography

The Crew Earth Observations (CEO) Facility within the Earth Science and Remote Sensing Unit at NASA’s Johnson Space Center supports the acquisition, analysis, and curation of astronaut photography of Earth’s surface and atmosphere. Astronauts on the International Space Station (ISS) respond to requests from CEO to acquire imagery of scientific and education targets, to include high profile targets in response to activations from the International Charter for Space & Major Disasters (also known as the International Disaster Charter, or IDC) and NASA’s Disasters Program. CEO facilitates the acquisition of astronaut photography in response to IDC events and delivers georeferenced data products to the United States Geological Survey (USGS) for distribution to the disaster community. Using GeoRef, an internal web-based tool developed in collaboration with NASA’s Ames Research Center, CEO generates data packages of georeferenced imagery, uncertainty images for assessing control and tie point accuracy, and metadata documenting raw and processed data. Operational experience with the Georef software identified vulnerabilities to internal code and server errors that can significantly increase time of data production. As such, CEO developed a backup procedure in case the GeoRef software experiences front-end or back-end errors. A system using OSGEO’s open-source QGIS software combined with a semi-automated pipeline using the object-oriented Python language and the Geospatial Abstract Library for generating metadata is quantitatively compared to GeoRef’s data package for quality and productivity. Root Mean Square Error (RMSE) provides a standard measurement of data quality as it relates to ground error. Assessing RMSE measurements generated from georeferenced astronaut photographs acquired with different obliquity and focal length offers a comprehensive accuracy assessment of the software’s transformation algorithms. This assessment will indicate the software's ability to produce data products with the least ground-error or highest data quality regarding ground accuracy. In addition, a comparison of the software’s efficiency in generating a data package that includes georeferenced images, metadata, and uncertainty images for measuring tie/ground point error was performed. Initial results, based on the comparison of three nadir-facing astronaut photographs acquired with a 95mm focal length, reveal the QGIS-based system's average RMSE is 2.36 (pixels) suggesting its georectification system produces data products that meet and perhaps improve upon Georef solution's average RMSE of 32.99 (pixels). However, the QGIS system was unable to reproduce two unique Georef data products, uncertainty images for measuring tie and control point errors and a translated unwrapped image. In addition, the Georef software is designed to accept handheld camera pose information from a hardware component (Geosens) scheduled for deployment on the ISS in late 2018; this information is intended to provide increased accuracy and auto-registration capability for astronaut photographs. Future work is expected to determine the QGIS-based georectification system’s potential as an open-source alternative (and operational backup) to Georef for georeferencing the full range of resolutions and viewing angles unique to handheld digital camera imagery in support of ISS disaster response activities.

Jagge, Amy M.↗

Intelligent Systems Technologies and Utilization of Earth Observation Data

The addition of raw data and derived geophysical parameters from several Earth observing satellites over the last decade to the data held by NASA data centers has created a data rich environment for the Earth science research and applications communities. The data products are being distributed to a large and diverse community of users. Due to advances in computational hardware, networks and communications, information management and software technologies, significant progress has been made in the last decade in archiving and providing data to users. However, to realize the full potential of the growing data archives, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system (KBS). Potential Intelligent Archive concepts include: 1) Mining archived data holdings to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services; 3) Recognizing the value of results, indexing and formatting them for easy access; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building; and 5) Being aware of other nodes in the KBS, participating in open systems interfaces and protocols for virtualization, and achieving collaborative interoperability.

Ramapriyan, H. K.↗

Evolving a NASA Digital Object Identifiers System with Community Engagement

To demonstrate how the ESDIS (Earth Science Data and Information System) DOI (Digital Object Identifier) system and its processes have evolved over these years based on the recommendations provided by the user community (whether the community members create and manage DOI information or use DOIs in the data citations). The user community is comprised of people with common interests and needs for data identifiers who are actively involved in the creation and usage process. Engagement describes the interactive context wherein the community provides information, evaluates the proposed processes, and provides guidance in the area of identifiers.

Identifiers↗

Standardizing Interfaces for External Access to Data and Processing for the NASA Ozone Product Evaluation and Test Element (PEATE)

NASA's traditional science data processing systems have focused on specific missions, and providing data access, processing and services to the funded science teams of those specific missions. Recently NASA has been modifying this stance, changing the focus from Missions to Measurements. Where a specific Mission has a discrete beginning and end, the Measurement considers long term data continuity across multiple missions. Total Column Ozone, a critical measurement of atmospheric composition, has been monitored for'decades on a series of Total Ozone Mapping Spectrometer (TOMS) instruments. Some important European missions also monitor ozone, including the Global Ozone Monitoring Experiment (GOME) and SCIAMACHY. With the U.S.IEuropean cooperative launch of the Dutch Ozone Monitoring Instrument (OMI) on NASA Aura satellite, and the GOME-2 instrumental on MetOp, the ozone monitoring record has been further extended. In conjunction with the U.S. Department of Defense (DoD) and the National Oceanic and Atmospheric Administration (NOAA), NASA is now preparing to evaluate data and algorithms for the next generation Ozone Mapping and Profiler Suite (OMPS) which will launch on the National Polar-orbiting Operational Environmental Satellite System (NPOESS) Preparatory Project (NPP) in 2010. NASA is constructing the Science Data Segment (SDS) which is comprised of several elements to evaluate the various NPP data products and algorithms. The NPP SDS Ozone Product Evaluation and Test Element (PEATE) will build on the heritage of the TOMS and OM1 mission based processing systems. The overall measurement based system that will encompass these efforts is the Atmospheric Composition Processing System (ACPS). We have extended the system to include access to publically available data sets from other instruments where feasible, including non-NASA missions as appropriate. The heritage system was largely monolithic providing a very controlled processing flow from data.ingest of satellite data to the ultimate archive of specific operational data products. The ACPS allows more open access with standard protocols including HTTP, SOAPIXML, RSS and various REST incarnations. External entities can be granted access to various modules within the system, including an extended data archive, metadata searching, production planning and processing. Data access is provided with very fine grained access control. It is possible to easily designate certain datasets as being available to the public, or restricted to groups of researchers, or limited strictly to the originator. This can be used, for example, to release one's best validated data to the public, but restrict the "new version" of data processed with a new, unproven algorithm until it is ready. Similarly, the system can provide access to algorithms, both as modifiable source code (where possible) and fully integrated executable Algorithm Plugin Packages (APPs). This enables researchers to download publically released versions of the processing algorithms and easily reproduce the processing remotely, while interacting with the ACPS. The algorithms can be modified allowing better experimentation and rapid improvement. The modified algorithms can be easily integrated back into the production system for large scale bulk processing to evaluate improvements. The system includes complete provenance tracking of algorithms, data and the entire processing environment. The origin of any data or algorithms is recorded and the entire history of the processing chains are stored such that a researcher can understand the entire data flow. Provenance is captured in a form suitable for the system to guarantee scientific reproducability of any data product it distributes even in cases where the physical data products themselves have been deleted due to space constraints. We are currently working on Semantic Web ontologies for representing the various provenance information. A new web site focusing on consolidating informaon about the measurement, processing system, and data access has been established to encourage interaction with the overall scientific community. We will describe the system, its data processing capabilities, and the methods the community can use to interact with the standard interfaces of the system.

Tilmes, Curt A.↗

SCDU Testbed Automated In-Situ Alignment, Data Acquisition and Analysis

In the course of fulfilling its mandate, the Spectral Calibration Development Unit (SCDU) testbed for SIM-Lite produces copious amounts of raw data. To effectively spend time attempting to understand the science driving the data, the team devised computerized automations to limit the time spent bringing the testbed to a healthy state and commanding it, and instead focus on analyzing the processed results. We developed a multi-layered scripting language that emphasized the scientific experiments we conducted, which drastically shortened our experiment scripts, improved their readability, and all-but-eliminated testbed operator errors. In addition to scientific experiment functions, we also developed a set of automated alignments that bring the testbed up to a well-aligned state with little more than the push of a button. These scripts were written in the scripting language, and in Matlab via an interface library, allowing all members of the team to augment the existing scripting language with complex analysis scripts. To keep track of these results, we created an easily-parseable state log in which we logged both the state of the testbed and relevant metadata. Finally, we designed a distributed processing system that allowed us to farm lengthy analyses to a collection of client computers which reported their results in a central log. Since these logs were parseable, we wrote query scripts that gave us an effortless way to compare results collected under different conditions. This paper serves as a case-study, detailing the motivating requirements for the decisions we made and explaining the implementation process.

Automation↗

An Update on the CDDIS

The Crustal Dynamics Data Inforn1ation System (CoorS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to store space geodesy and geodynamics related data products in a central data bank, to maintain infom1ation about the archival of these data, and to disseminate these data and information in a timely mam1er to a global scientific research community. The archive consists of GNSS, laser ranging, VLBI, and OORIS data sets and products derived from these data. The coors is one of NASA's Earth Observing System Oata and Infom1ation System (EOSorS) distributed data centers; EOSOIS data centers serve a diverse user community and are tasked to provide facilities to search and access science data and products. The coors data system and its archive have become increasingly important to many national and international science communities, in pal1icular several of the operational services within the International Association of Geodesy (lAG) and its project the Global Geodetic Observing System (GGOS), including the International OORIS Service (IDS), the International GNSS Service (IGS), the International Laser Ranging Service (ILRS), the International VLBI Service for Geodesy and Astrometry (IVS), and the International Earth Rotation Service (IERS). The coors has recently expanded its archive to supp011 the IGS Multi-GNSS Experiment (MGEX). The archive now contains daily and hourly 3D-second and subhourly I-second data from an additional 35+ stations in RINEX V3 fOm1at. The coors will soon install an Ntrip broadcast relay to support the activities of the IGS Real-Time Pilot Project (RTPP) and the future Real-Time IGS Service. The coors has also developed a new web-based application to aid users in data discovery, both within the current community and beyond. To enable this data discovery application, the CDDIS is currently implementing modifications to the metadata extracted from incoming data and product files pushed to its archive. This poster will include background information about the system and its user communities, archive contents and updates, enhancements for data discovery, new system architecture, and future plans.

Noll, Carey↗

Marsviewer

Marsviewer is a multi-platform application designed to aid in quality control, browsing, and analysis of original science product images (Experiment Data Records, or EDRs) and derived image data products (Reduced Data Records, or RDRs) returned by the Mars Explorer Rover (MER) mission. Marsviewer offers an abstraction of the products organization via a file finder. For example, the application understands the file structure and filename conventions of the MER Operational Storage Server, helping the user to navigate this complex file system to find desired images. Marsviewer also works with a flat file system, remote-operations file systems, image-archive file systems, and others. All EDRs found for a given solar day (Sol) are displayed in a list, optionally with thumbnail images. Once the user selects an image from the list, a tabbed pane conveniently displays the original source image and all associated RDRs. Marsviewer provides the option of overlaying derived images upon the source image, resulting in an easier-to-interpret color representation of the data. Display manipulations such as zoom, data range adjustment, contrast enhancement, and contour control are available. Image metadata (labels) from the current image can be displayed and searched. The architecture of the program is extensible: new types of RDRs can be installed and new file finders can be added to adapt the program to different file structures and different filename conventions. This keeps the application flexible and provides an opportunity for reuse with future rover missions.

Toole, Nicholas↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗

Earth Science Datacasting v2.0

The Datacasting software, which consists of a server and a client, has been developed as part of the Earth Science (ES) Datacasting project. The goal of ES Datacasting is to provide scientists the ability to automatically and continuously download Earth science data that meets a precise, predefined need, and then to instantaneously visualize it on a local computer. This is achieved by applying the concept of podcasting to deliver science data over the Internet using RSS (Really Simple Syndication) XML feeds. By extending the RSS specification, scientists can filter a feed and only download the files that are required for a particular application (for example, only files that contain information about a particular event, such as a hurricane or flood). The extension also provides the ability for the client to understand the format of the data and visualize the information locally. The server part enables a data provider to create and serve basic Datacasting (RSS-based) feeds. The user can subscribe to any number of feeds, view the information related to each item contained within a feed (including browse pre-made images), manually download files associated with items, and place these files in a local store. The client-server architecture enables users to: a) Subscribe and interpret multiple Datacasting feeds (same look and feel as a typical mail client), b) Maintain a list of all items within each feed, c) Enable filtering on the lists based on different metadata attributes contained within the feed (list will reference only data files of interest), d) Visualize the reference data and associated metadata, e) Download files referenced within the list, and f) Automatically download files as new items become available.

Bingham, Andrew W.↗

Updates of Land Surface and Air Quality Products in NASA MAIRS and NEESPI Data Portals

Following successful support of the Northern Eurasia Earth Sciences Partner Initiative (NEESPI) project with NASA satellite remote sensing data, from Spring 2009 the NASA GES DISC (Goddard Earth Sciences Data and Information Services Center) has been working on collecting more satellite and model data to support the Monsoon Asia Integrated Regional Study (MAIRS) project. The established data management and service infrastructure developed for NEESPI has been used and improved for MAIRS support.Data search, subsetting, and download functions are available through a single system. A customized Giovanni system has been created for MAIRS.The Web-based on line data analysis and visualization system, Giovanni (Goddard Interactive Online Visualization ANd aNalysis Infrastructure) allows scientists to explore, quickly analyze, and download data easily without learning the original data structure and format. Giovanni MAIRS includes satellite observations from multiple sensors and model output from the NASA Global Land Data Assimilation System (GLDAS), and from the NASA atmospheric reanalysis project, MERRA. Currently, we are working on processing and integrating higher resolution land data in to Giovanni, such as vegetation index, land surface temperature, and active fire at 5km or 1km from the standard MODIS products. For data that are not archived at the GESDISC,a product metadata portal is under development to serve as a gateway for providing product level information and data access links, which include both satellite, model products and ground-based measurements information collected from MAIRS scientists.Due to the large overlap of geographic coverage and many similar scientific interests of NEESPI and MAIRS, these data and tools will serve both projects.

Shen, Suhung↗