Search NASA⌕ Search

NASA NTRS · 20170000324

Hadoop for High-Performance Climate Analytics: Use Cases and Lessons Learned

Abstract

Scientific data services are a critical aspect of the NASA Center for Climate Simulations mission (NCCS). Hadoop, via MapReduce, provides an approach to high-performance analytics that is proving to be useful to data intensive problems in climate research. It offers an analysis paradigm that uses clusters of computers and combines distributed storage of large data sets with parallel computation. The NCCS is particularly interested in the potential of Hadoop to speed up basic operations common to a wide range of analyses. In order to evaluate this potential, we prototyped a series of canonical MapReduce operations over a test suite of observational and climate simulation datasets. The initial focus was on averaging operations over arbitrary spatial and temporal extents within Modern Era Retrospective- Analysis for Research and Applications (MERRA) data. After preliminary results suggested that this approach improves efficiencies within data intensive analytic workflows, we invested in building a cyber infrastructure resource for developing a new generation of climate data analysis capabilities using Hadoop. This resource is focused on reducing the time spent in the preparation of reanalysis data used in data-model inter-comparison, a long sought goal of the climate community. This paper summarizes the related use cases and lessons learned.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tamkin, Glenn. 2013-06-26. Hadoop for High-Performance Climate Analytics: Use Cases and Lessons Learned. https://ntrs.nasa.gov/citations/20170000324

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

SeeQ: A Programming Model for Portable Data-Driven Building Applications

This paper introduces SeeQ, a programming model and an abstraction framework that facilitates the development of portable data- driven building applications. Data-driven approaches can provide insights into building operations and guide decision-making to achieve operational objectives. Yet the configuration of such applications per building requires extensive effort and tacit knowledge. In SeeQ, we propose a portable programming model and build a software system that enables self-configuration and execution across diverse buildings. The configuration of each building is captured in a unified data model - in this paper, we work with the Brick ontology without loss of generality. SeeQ focuses on the distinction between the application logic and the configuration of an application against building-specific data inputs and systems. We test the proposed approach by configuring and deploying a diverse range of applications across five heterogeneous real-world buildings. The analysis shows the potential of SeeQ to significantly reduce the efforts associated with the delivery of building analytics.

analytics↗

Low Carbon Technology Strategies: Large Office

This document includes steps that building owners and operators can implement to achieve smart, healthy, and low-carbon large office buildings within their existing building portfolios. Large offices are typically over 50,000 square feet and often include complex heating and cooling systems.

analytics↗

Low Carbon Technology Strategies: Small Office

This document includes steps that building owners and operators can implement to achieve smart, healthy, and low-carbon small office buildings within their existing building portfolios. Small offices are typically less than 50,000 square feet and often use packaged rooftop units for heating, cooling, and ventilation.

analytics↗