Search NASA⌕ Search

DOE OSTI · 2438552

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

Abstract

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore’s Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to MLexperts because of lack of training data and relevant, easy-to-use workflows. Wepresent ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state. The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res2) and code(https://leap-stc.github.io/ClimSim)arereleasedopenlytosupport the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yu, Sungduk, Hannah, Walter M., Peng, Liran, Lin, Jerry, Bhouri, Mohamed A., Gupta, Ritwik, Lutjens, Bjorn, Will, Justus C., Behrens, Gunnar, Busecke, Julius J., Loose, Nora, Stern, Charles, Beucler, Tom, Harrop, Bryce E., Hillman, Benjamin R., Jenney, Andrea M., Ferretti, Savannah L., Liu, Nana, Anandkumar, Anima, Brenowitz, Noah, Eyring, Veronika, Geneva, Nicholas, Gentine, Pierre, Mandt, Stephan, Pathak, Jaideep, Subramaniam, Akshay, Vondrick, Carl, YU, ROSE, Zanna, Laure, Zheng, Tian, Abernathey, Ryan P., Ahmed, Fiaz, Bader, David C., Baldi, Pierre, Barnes, Elizabeth, Bretherton, Christopher S., Caldwell, Peter M., Chuang, Wayne, Han, Yilun, Huang, Yu, Iglesias-Suarez, Fernando, Jantre, Sanket, Kashinath, Karthik, Khairoutdinov, Marat, Kurth, Thorsten, Lutsko, Nicholas J., Ma, Po Lun, Mooers, Griffin, Neelin, David, Randall, David A., Shamekh, Sara, Taylor, Mark, Urban, Nathan M., Yuval, Janni, Zhang, Guang J., Pritchard, Michael S.. 2023-12-29. ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation. https://www.osti.gov/biblio/2438552

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Physics-informed Deep Reinforcement Learning-based Control in Power systems

Incorporating physics information into the deep reinforcement learning (DRL) process is a promising approach for addressing the challenges faced in learning-based control design problems for physical systems. Power grid dynamics, being a physical system, adheres to specific physical laws, constraints, as well as operational and control rules. Therefore, consideration of such physics-based law improves the learning process drastically. In general, traditional grid control schemes rely on rule-based mechanisms that cannot adapt to changing operating conditions. To improve the adaptability and computation time, recent research has seen a surge of DRL-based applications in power grid control. A generic DRL-based control design imposes the system performance requirements through the design of reward functions. In some cases, some of the important physics information is injected through this reward function. However, due to the complex dynamics and large state-action space, learning an optimal DRL policy often becomes challenging. Inspired by the latest developments in general machine learning (ML) research, power system researchers have been investigating more direct ways of incorporating physics knowledge into DRL training. This chapter specifically focuses on these aspects of physics-informed DRL designs in grid control. It discusses the significance, applications, research gaps, and open problems that need to be addressed in future research.

artificial intelligence, machine learning↗