Search NASA⌕ Search

SEARCH · Search NASA

Results for “Google Engine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Investigating the Impacts of Land Use Change on Urban Heat and Vulnerability in Cali, Columbia

The urban heat island effect (UHI) is an environmental phenomenon where cities experience higher temperatures than rural areas due to increased pavement and decreased cooling from vegetation. Approximately 76% of people in Colombia live in urban areas, and the city of Santiago de Cali is facing UHI challenges exacerbated by land use change. Wetlands and forests formerly surrounded the city but were replaced by development and agriculture. The Colombian municipal government agency Departamento Administrativo de Gestión del Medio Ambiente and the community organization Fundacion Dinamizadores Ambientales partnered with NASA DEVELOP to evaluate communities in Cali most vulnerable to urban heat. This project illustrated the utility of using NASA Earth observations to evaluate the relationship between land use, temperature, and social factors in Cali, Colombia between 2013 and 2023. The team used Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI) and Thermal Infrared Sensor (TIRS), and Landsat 9 OLI-2/TIRS-2 to generate land surface temperature (LST), normalized difference vegetation index (NDVI), and albedo maps in Google Earth Engine. Heavy cloud cover limited the accuracy of the LST but incorporating up to three satellites for a median image reduced potential errors. Through further analysis in ArcGIS Pro, the team classified land use change using a deep learning model and found that LST was significantly higher in urban areas than in wetlands or forests. Using R studio, the team ran a principal component analysis to determine which social factors had the strongest correlation with LST. The team found that health care and green space access were negatively correlated, and Afro-Colombian ethnicity was positively correlated with LST. With awareness of the most impacted and vulnerable regions, the partner organizations can work to prioritize green space establishment in those areas to reduce the impacts of urban heat. Addressing the urban heat island effect will reduce environmental justice concerns within the city and improve overall health, air, and water quality for those who live there.

Brenna Bruffey↗

Utilizing Earth Observations to Understand Landscape Patterns and Assist in Wildlife Management in Iona National Park, Angola

Following the end of the Angolan Civil War (1975-2002), human habitation in Iona National Park has grown exponentially, as has the livestock population. An ongoing drought beginning in 2017 has brought people, livestock, and wildlife into increasing competition for resources within the park. This study used Earth observation data, primarily Landsat and Sentinel imagery, to examine landscape trends to improve wildlife preservation approaches in Iona National Park, Angola. In collaboration with the NGO African Parks, we developed a robust land use and land cover (LULC) classification model using remote sensing data to augment sparse ground-based data in this arid land region. We used Google Earth Engine and a random forest classifier to map vegetation types, water bodies, and potential wildlife habitats. This analysis resulted in a high spatial resolution LULC time-series between 1984-2023, highlighting critical periods of socioecological change over the past 40 years. These results increased the partner’s ability to make scientifically grounded decisions about resource allocation and conservation priorities. This analysis supports the feasibility of applying remote sensing techniques coupled with machine learning models in dry regions, where standard survey methods are frequently limited by accessibility and resource availability. However, we identified limitations in ground-truth data and the difficulty of recognizing certain vegetation types in arid areas. Despite these limitations, the study demonstrated Earth observations' ability to transform wildlife management techniques in distant and data-scarce locations, providing a reproducible foundation for similar ecosystems around the world.

Emmanuel Aklie↗

Cloud-Based Solutions for Monitoring Coastal Ecosystems and Prioritization of Restoration Efforts Across Belize

In recent years the availability of automated change detection algorithms in Google Earth Engine has permitted cloud-based processing of large time series of satellite imagery. Models such as the Continuous Change Detection and Classification (CCDC), CCDC-Spectral Mixture Analysis (CCDC-SMA), and Landsat-based Detection of Trends in Disturbance and Recovery (LandTrendr) allow users to exploit decades of Earth Observations (EO) , leveraging the Landsat archive and data from other sensors to detect disturbance in forest ecosystems. Despite the wide adoption of these methods, robust documentation and growing community of users, little research has explored their use in mangrove environments. Mangroves are dynamic environments subject to changes due to not only the natural migration of mudflats but also coastal erosion, urban expansion, aquaculture practices, etc. This work aims to identify best practices for the application of these models to identify and monitor changes in the Belizean mangroves, which experienced an estimated 5.4% decrease in extent between 1980 and 2017 (Cherrington et al. 2020). Partnering directly with the Belizean Forest Department, our team will develop a replicable, efficient methodology to annually update the country’s mangrove extent employing EO-based change detection. This collaboration supports Belize’s marine conservation and climate resilience efforts, such as its Blue Bond agreement with the Nature Conservancy, building on national processes to monitor these critical ecosystems.

Mangroves↗

Increasing Discovery and Usability of Earth Science Satellite Data with My NASA Data

For 20 years, the My NASA Data project at NASA Langley Research Center has developed innovative approaches to increase the use of NASA’s satellite data by learners. My NASA Data offers a variety of authentic Earth Science datasets and a data visualization tool, eliminating the need for educators and/or learners to obtain specialized knowledge of GIS data formats and software to access and use authentic Earth Science data. While there is no shortage of available data, as federal government agencies such as NASA house petabytes of freely accessible Earth Science datasets, much of the data are only available for download and visualization in specialized formats and software, limiting their accessibility to educators and learners, especially those in primary and secondary school. Using the Google Earth Engine platform, the My NASA Data team has recently reinvented their data visualization tool, called the Earth System Data Explorer (ESDE). The ESDE gives users the capability to explore over 60 Earth Science satellite datasets in a multitude of formats such as maps, graphs, and data table Its new and improved user interface design was developed based on the preferences of educators, whom the My NASA Data project has over 20 years’ experience working with. Earth Science and GIS Subject Matter Experts (SMEs) structured the data in a professional and scientific manner. During Fiscal Year 2023, the My NASA Data website received over 1 million digital engagements, with over one-third being visitors to the data visualization tool. These metrics highlight the interest in a visualization tool that is simple and free to use with reliable and trusted datasets. The ESDE empowers users to readily relate and analyze NASA Earth Science data within their area of interest. The team used a user-centered design (UCD) framework to receive and incorporate feedback into the application’s design. Core requested features include the ability to create time series graphs, comparative analysis of maps, and download the data as CSV file. Responses indicate that advances in data visualization tools such as the ESDE make authentic Earth Science data more accessible. This presentation will cover how the My NASA Data project develops tools to enhance data discovery and accessibility, as well as how SME and user suggestions are incorporated.

Desiray Wilson↗

Smart Hand Offs and Earthdata Search

Often a science user will discover data of interest in a general-purpose discovery tool like Earthdata Search. At that point it might be beneficial for the science user to switch to a more specialized tool. Smart handoffs allow the user to switch from one tool to another with their context (collection, spatial and temporal constraints) intact. This provides an efficient means of traversing tools and services. The idea of smart handoffs could be expanded to commercial search engines such as Google, visualization tools like State of the Ocean and Giovanni and domain-specific search tools like Earthdata Search.

Giovanni↗

Pacific Northwest Health & Air Quality - Utilizing NASA Earth Observations to Analyze Air Quality Impacts from Wildfires in the Pacific Northwest

The Pacific Northwest region of the United States and Canada has become more vulnerable to intense wildfire regimes due to years of fire suppression and climatic changes. Smoke from fires exposes communities to hazardous aerosols and pollutants known to trigger asthma symptoms and exacerbate other respiratory and cardiovascular diseases. In partnership with The Nature Conservancy’s Washington Chapter and the Puget Sound Clean Air Agency, NASA DEVELOP investigated the impacts of wildfire smoke on air quality from 2008 to 2020 using NASA Earth observations. To explore the various dimensions of smoke and its relation to air quality, the team looked at the vertical extent of smoke plumes and the resulting changes in air quality. The team evaluated the potential relationship between plume height of wildfire smoke and fire radiative power using the Moderate Resolution Imaging Spectroradiometer (MODIS) aboard the Aqua and Terra satellites and Terra’s Multi-angle Imaging SpectroRadiometer (MISR) using the MISR INteractive eXplorer (MINX). The team determined that there was no regional relationship between fire radiative power and smoke plume height. To investigate changes in air quality resulting from wildfire smoke, the team utilized data from NASA’s Fire Information for Resource Management System, from the European Space Agency’s Sentinel-5P Tropospheric Monitoring Instrument (TROPOMI), and true color imagery from Landsat 8 Operational Land Imager (OLI). The team created a Google Earth Engine-based (GEE) web tool to visualize changes in atmospheric pollutants and aerosol optical depth. Results of case study fires showed varying increases in pollutant concentrations when compared to a baseline map. The end products provided the partners with tools to quantify plume height using MINX and to visualize recent air quality patterns relating to variations in wildfire extent and severity in the Pacific Northwest

Ani Matevosian↗

Automated Collection of Scientific Publications Linked to NASA Earth Science Datasets

NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.

Infometrics↗

Interplay of Topography, Fire History, and Climate on Interior Alaska Boreal Forest Vegetation Dynamics in the 21st Century: A Landsat Time-Series Analysis

This study investigates vegetation dynamics in boreal forests of Interior Alaska, focusing on topography, fire history, and climate influences. The study area includes Bonanza Creek Experimental Forest (BCEF) and surrounding region, categorized by topography (upland, floodplain, lowland) and fire history. Using Mann–Kendall trend and Theil–Sen slope analyses on Landsat-derived spectral metrics: Normalized Difference Vegetation Index (NDVI) and Normalized Burn Ratio (NBR), we observed a shift from browning to greening trends, particularly in historically burned areas. The photosynthetic activity in burned upland converged with unburned areas ~30 years post-fire, coincident with a shift towards deciduous dominance during post-fire succession. Normalized Difference Moisture Index (NDMI) trends revealed a significant increase in vegetation moisture content across all topographies. We introduce Effective Seasonal Precipitation Index (ESPI), which combines prior-year annual precipitation with current-year spring snow depth. Its positive correlation with NDMI highlights its potential for monitoring vegetation moisture dynamics at the landscape scale. Furthermore, by correlating dendrochronology-based climate indices, we found strong correlation between NDMI and normalized Supplemental Precipitation Index (nSPI), across topographies. Overall, this research provides critical insights into how climate and fire influence interior boreal vegetation, highlighting the effects of increased precipitation, and topography on shaping differential vegetation responses across the landscape.

Google Earth Engine↗

Using NASA's Giovanni Web Portal to Access and Visualize Satellite-based Earth Science Data in the Classroom

One of the biggest obstacles for the average Earth science student today is locating and obtaining satellite-based remote sensing data sets in a format that is accessible and optimal for their data analysis needs. At the Goddard Earth Sciences Data and Information Services Center (GES-DISC) alone, on the order of hundreds of Terabytes of data are available for distribution to scientists, students and the general public. The single biggest and time-consuming hurdle for most students when they begin their study of the various datasets is how to slog through this mountain of data to arrive at a properly sub-setted and manageable data set to answer their science question(s). The GES DISC provides a number of tools for data access and visualization, including the Google-like Mirador search engine and the powerful GES-DISC Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni) web interface.

Lloyd, Steven↗

Generic, Extensible, Configurable Push-Pull Framework for Large-Scale Science Missions

The push-pull framework was developed in hopes that an infrastructure would be created that could literally connect to any given remote site, and (given a set of restrictions) download files from that remote site based on those restrictions. The Cataloging and Archiving Service (CAS) has recently been re-architected and re-factored in its canonical services, including file management, workflow management, and resource management. Additionally, a generic CAS Crawling Framework was built based on motivation from Apache s open-source search engine project called Nutch. Nutch is an Apache effort to provide search engine services (akin to Google), including crawling, parsing, content analysis, and indexing. It has produced several stable software releases, and is currently used in production services at companies such as Yahoo, and at NASA's Planetary Data System. The CAS Crawling Framework supports many of the Nutch Crawler's generic services, including metadata extraction, crawling, and ingestion. However, one service that was not ported over from Nutch is a generic protocol layer service that allows the Nutch crawler to obtain content using protocol plug-ins that download content using implementations of remote protocols, such as HTTP, FTP, WinNT file system, HTTPS, etc. Such a generic protocol layer would greatly aid in the CAS Crawling Framework, as the layer would allow the framework to generically obtain content (i.e., data products) from remote sites using protocols such as FTP and others. Augmented with this capability, the Orbiting Carbon Observatory (OCO) and NPP (NPOESS Preparatory Project) Sounder PEATE (Product Evaluation and Analysis Tools Elements) would be provided with an infrastructure to support generic FTP-based pull access to remote data products, obviating the need for any specialized software outside of the context of their existing process control systems. This extensible configurable framework was created in Java, and allows the use of different underlying communication middleware (at present, both XMLRPC, and RMI). In addition, the framework is entirely suitable in a multi-mission environment and is supporting both NPP Sounder PEATE and the OCO Mission. Both systems involve tasks such as high-throughput job processing, terabyte-scale data management, and science computing facilities. NPP Sounder PEATE is already using the push-pull framework to accept hundreds of gigabytes of IASI (infrared atmospheric sounding interferometer) data, and is in preparation to accept CRIMS (Cross-track Infrared Microwave Sounding Suite) data. OCO will leverage the framework to download MODIS, CloudSat, and other ancillary data products for use in the high-performance Level 2 Science Algorithm. The National Cancer Institute is also evaluating the framework for use in sharing and disseminating cancer research data through its Early Detection Research Network (EDRN).

Foster, Brian M.↗

Local Scale (3-M) Soil Moisture Mapping Using SMAP and Planet Superdove

A capability for mapping meter-level resolution soil moisture with frequent temporal sampling over large regions is essential for quantifying local-scale environmental heterogeneity and eco-hydrologic behavior. However, available surface soil moisture (SSM) products generally involve much coarser grain sizes ranging from 30 m to several 10s of kilometers. Hence a new method is proposed to estimate 3-m resolution SSM using a combination of multi-sensor fusion, machine- learning (ML) and Cumulative Distribution Function (CDF) matching approaches. This method established favorable SSM correspondence between 3-m pixels and overlying 9-km grid cells from overlapping Planet SuperDove (PSD) observations and NASA Soil Moisture Active-Passive (SMAP) mission products. The resulting 3-m SSM predictions showed improved accuracy by reducing ab- solute bias and RMSE by ~0.01 cm3/cm3 over the original SMAP data in relation to in-situ soil moisture measurements for the Australian Yanco region, while preserving the high sampling frequency (1-3 day global revisit) and sensitivity to surface wetness (R 0.865) from SMAP. Heterogeneous soil moisture distributions varying with vegetation biomass gradients and irrigation regimes were generally captured within a selected study area. Further algorithm refinement and implementation for regional applications will allow for improvement in water resources management, precision agriculture, and disaster forecasts and responses.

soil moisture↗

Front Range Wildland Fires: Evaluating the Efficacy of Remote Sensing Imagery in Monitoring Forest Fuels Treatment Methods

Over the last several decades, wildfire frequency and severity in forested areas along Colorado’s Front Range have increased due to a buildup of fuels. This has led to an increase in forest treatments, as well as an increased need to evaluate the success of these treatments. Remote sensing products offer an efficient and cost-effective way to monitor forest treatments; however, not all remote sensing products and analysis techniques have been explored by Coloradan land managers. Specifically, project partners at the Colorado State Forest Service (CSFS) and the Colorado Forest Restoration Institute (CFRI) were interested in using an effective and streamlined method of mapping canopy cover to better monitor forest treatment success. To support their needs, the NASA DEVELOP Front Range Wildland Fires team explored National Agricultural Imagery Program (NAIP) imagery at different spatial resolutions and numbers of training points with NASA’s Shuttle Radar Topography Mission (SRTM) Data Elevation Model (DEM) as a predictor in addition to NAIP imagery spectral predictors. From this analysis, we created classified canopy cover rasters, and compared accuracy metrics across model iterations. We also determined that the best performing model, with an overall accuracy of 0.900 uses 2021 NAIP imagery at 2-meter resolution, 800 training points, 200 testing points, does not use topographic predictors, and reclassifies shadow pixels via a pre-selected NDVI threshold.

Remote Sensing↗

Great Salt Lake Health and Air Quality: Monitoring Lakebed Exposure and its Impact on Air Quality and Environmental Hazards in the Great Salt Lake Watershed

Water flow into the Great Salt Lake has declined rapidly over the last forty years due to human withdrawals and climate change. As a result of declining lake levels, over 50% of the lakebed is now exposed. Dust storms may grow in frequency and intensity across Northern Utah as lakebed dust becomes airborne under specific meteorological conditions. In our research project, we utilized satellite imagery from Terra and Aqua, Sentinel-5P, CALIPSO, Landsat 5 TM, Landsat 7 ETM+, Landsat 8 OLI-2, Suomi NPP, ground sensor environmental data, and demographic data to understand the relationship between lake desiccation and dust, and the impact of pollution upon the communities surrounding the Great Salt Lake. By plotting changes in Lake Surface Area against Aerosol Optical Depth (AOD) over our study period (2010-2022), we found an inverse relationship (R2=0.3423) between lake surface area and dust levels within our study area. We conducted a Vertical Feature Mask (VFM) and Extinction Coefficient Plot, from which we identified that during dust events, the aerosol type is mainly polluted dust and the aerosol height is 200 meters from the surface. Lastly, we created bivariate choropleth maps, which demonstrate which census tracts within our study area are most vulnerable to AOD (a proxy for PM2.5 from dust), NO2 and HCHO (precursors to ozone). In summary, our findings revealed that declining lake levels are associated with an increase in intensity of dust events, and these dust events will particularly impact residents of Tooele County and the west side of Salt Lake City. Project resources support partner needs by informing targeted air monitoring efforts, lakebed management practices, and advocacy efforts for GSL stewardship.

Terminal Saline Lake↗

Lake Anna Water Resources: Using NASA Earth Observations to Identify Algal Event Risk Factors in Lake Anna and Help Inform Future Management Practices

Lake Anna is a man-made reservoir and popular recreation destination that spans over 13,000 acres-9,600 public and 3,400 private-in the Piedmont region of Virginia. The Lake has recently seen a rise in documented harmful algal blooms (HABs), which pose a variety of community and ecological concerns and are often exacerbated by anthropogenic factors, such as excess nutrient loads from agricultural runoff. NASA DEVELOP has partnered with the Virginia Department of Environmental Quality (DEQ) to help monitor cyanobacteria and nutrient pollution indicators across Lake Anna. The team utilized Earth observations (EO) and in situ ancillary data to identify and monitor algal bloom trends. The team used Landsat 8 Operational Land Imager (OLI), Landsat 9 OLI-2, Sentinel-2 Multispectral Instrument (MSI), and Sentinel-3 Ocean and Land Color Instrument (OLCI) to analyze water quality variables such as chlorophyll-a, turbidity, surface temperature, and cyanobacteria. Due to the lack of comprehensive in situ data and historic HAB event records, our ability to compare and validate EOs was limited. Additionally, deficient spatial resolutions, along with a geographically complicated shoreline, accentuated the spatial constraints we faced in our analysis. After examining EOs, our results indicated conditions conducive to the formation of HABs within the upper reaches of Lake Anna. Yet, spatially dependent limiting factors may have also influenced where these phenomena developed. When used in concert with existing in situ datasets, NASA EOs provide relevant stakeholders with more comprehensive analyses with which to engage in enhanced monitoring and watershed management.

cyanobacteria↗

A Machine-Learning Approach to Assess Aircraft Engine System Performance

Artificial intelligence (AI)/machine learning, and big data are transforming the global business environment. They have become the most disruptive technologies for organizations to improve workplace efficiency and productivity. This work explored the application of machine learning-based predictive analytics that would enable aircraft engine designers to estimate engine system performance quickly during the conceptual design stage. Supervised machine-learning algorithm was employed to study patterns in an existing database of production and research turbofan engines, and built predictive analytics for use in predicting system performance of new turbofan designs. Specifically, the author developed deep-learning analytics to predict turbofan system weight, using turbofan design parameters as the input. The predictive analytics were trained and deployed in Keras, an open-source neural networks API (application program interface) written in Python, with TensorFlow (an open-source artificial AI library developed by Google) serving as the backend engine. The current engine-weight prediction results, together with those for the TSFC (thrust specific fuel consumption) and core-size predictions that were studied previously by the author, show that machine learning-based predictive analytics can be an effective, time-saving tool for aircraft engine design-space exploration during the conceptual design stage. It would enable expeditious identification of the best engine design amongst several candidates.

Michael T Tong↗

Improved Search Techniques

Thousands of millions of documents are stored and updated daily in the World Wide Web. Most of the information is not efficiently organized to build knowledge from the stored data. Nowadays, search engines are mainly used by users who rely on their skills to look for the information needed. This paper presents different techniques search engine users can apply in Google Search to improve the relevancy of search results. According to the Pew Research Center, the average person spends eight hours a month searching for the right information. For instance, a company that employs 1000 employees wastes $2.5 million dollars on looking for nonexistent and/or not found information. The cost is very high because decisions are made based on the information that is readily available to use. Whenever the information necessary to formulate an argument is not available or found, poor decisions may be made and mistakes will be more likely to occur. Also, the survey indicates that only 56% of Google users feel confident with their current search skills. Moreover, just 76% of the information that is available on the Internet is accurate.

Albornoz, Caleb Ronald↗

Community-Based Services that Facilitate Interoperability and Intercomparison of Precipitation Datasets from Multiple Sources

Over the past 12 years, large volumes of precipitation data have been generated from space-based observatories (e.g., TRMM), merging of data products (e.g., gridded 3B42), models (e.g., GMAO), climatologies (e.g., Chang SSM/I derived rain indices), field campaigns, and ground-based measuring stations. The science research, applications, and education communities have greatly benefited from the unrestricted availability of these data from the Goddard Earth Sciences Data and Information Services Center (GES DISC) and, in particular, the services tailored toward precipitation data access and usability. In addition, tools and services that are responsive to the expressed evolving needs of the precipitation data user communities have been developed at the Precipitation Data and Information Services Center (PDISC) (http://disc.gsfc.nasa.gov/precipitation or google NASA PDISC), located at the GES DISC, to provide users with quick data exploration and access capabilities. In recent years, data management and access services have become increasingly sophisticated, such that they now afford researchers, particularly those interested in multi-data set science analysis and/or data validation, the ability to homogenize data sets, in order to apply multi-variant, comparison, and evaluation functions. Included in these services is the ability to capture data quality and data provenance. These interoperability services can be directly applied to future data sets, such as those from the Global Precipitation Measurement (GPM) mission. This presentation describes the data sets and services at the PDISC that are currently used by precipitation science and applications researchers, and which will be enhanced in preparation for GPM and associated multi-sensor data research. Specifically, the GES-DISC Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni) will be illustrated. Giovanni enables scientific exploration of Earth science data without researchers having to perform the complicated data access and match-up processes. In addition, PDISC tool and service capabilities being adapted for GPM data will be described, including the Google-like Mirador data search and access engine; semantic technology to help manage large amounts of multi-sensor data and their relationships; data access through various Web services (e.g., OPeNDAP, GDS, WMS, WCS); conversion to various formats (e.g., netCDF, HDF, KML (for Google Earth)); visualization and analysis of Level 2 data profiles and maps; parameter and spatial subsetting; time and temporal aggregation; regridding; data version control and provenance; continuous archive verification; and expertise in data-related standards and interoperability. The goal of providing these services is to further the progress towards a common framework by which data analysis/validation can be more easily accomplished.

Liu, Zhong↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗