Search NASA⌕ Search

DOE OSTI · 1841140

Two-Stage Reinforcement Learning Policy Search for Grid-Interactive Building Control

Abstract

This paper develops an intelligent grid-interactive building controller, which optimizes building operation during both normal hours and demand response (DR) events. To avoid costly on-demand computation and to adapt to non-linear building models, the controller utilizes reinforcement learning (RL) and makes real-time decisions based on a near-optimal control policy. Learning such a policy typically amounts to solving a hard non-convex optimization problem. We propose to address this problem with a novel global-local policy search method. In the first stage, an RL algorithm based on zero-order gradient estimation is leveraged to search for the optimal policy globally, due to its scalability and the potential to escape some poor performing local optima. The obtained policy is then fine-tuned locally to bring the first-stage solution closer to that of the original unsmoothed problem. Experiments on a simulated five-zone commercial building demonstrate the advantages of the proposed method over existing learning approaches. They also show that the learned control policy outperforms a pragmatic linear model predictive controller (MPC) and approaches the performance of an oracle MPC in testing scenarios. Using a state-of-the-art advanced computing system, we demonstrate that the controller can be learned and deployed within hours of training.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhang, Xiangyu, Chen, Yue, Bernstein, Andrey, Chintala, Rohit, Graf, Peter, Jin, Xin, Biagioni, David. 2022-01-10. Two-Stage Reinforcement Learning Policy Search for Grid-Interactive Building Control. https://doi.org/10.1109/tsg.2022.3141625

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Building the Business Case with JUSTIFI: Quantifying Multiple Benefits in the Automotive Industry

The Michigan State University Industrial Training and Assessment Center (MSU ITAC) conducted a pilot study at an automotive parts manufacturer in Michigan. The study identified energy-productivity enhancements through the application of the JUSTIFI software. Key recommendations included replacing six inefficient rooftop units (RTUs) with a new air rotational unit. By quantifying operational savings for this project, the expected payback period went from 6.6 years to 1.4 years. Additionally, the installation of variable frequency drives (VFDs) on condenser tower motors was suggested. By including all operational benefits, the payback period was reduced from 8.2 years to 1.3 years. This comprehensive analysis aims to bolster the manufacturer's goals of reducing energy while enhancing overall operational efficiency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Field Validation of a Grid-Interactive Efficient Building Software Solution

The U.S. General Services Administration's (GSA's) Green Proving Ground (GPG) program, in partnership with the National Laboratory of the Rockies (NLR), completed a field study of a Grid-Interactive Efficient Buildings (GEB) software solution. The study focused on a single testbed facility to test the GEB functionality of the software solution, along with other features. The testbed facility - a courthouse - is a common building type in GSA's vast building portfolio, offering potentially impactful findings on a scalable level. The study evaluated Prescriptive Data's technology, Nantum OS, a connected building operating system ("GEB Solution") which aggregates multiple sources of previously siloed building data and combines that data with external sources, such as weather information or utility signals, into a single integrated platform. A GEB Solution is a type of Energy Management Information System (EMIS). EMIS is defined as a system of devices, data services, and software applications that communicates with any building system or third-party data source to aggregate and transform data into new capabilities to aid in the optimization of energy use at the building, campus, or agency level. This specific GEB Solution is an EMIS with ASO, automated system optimization, offering supervisory control of certain aspects of the Building Automation System (BAS). Multiple features were evaluated including, but not limited to, Continuous Demand Management to avoid setting new monthly kilowatt (kW) peaks, energy efficiency for reduction of kilowatt hours (kWh) and natural gas consumption, and automated demand response (ADR) for purposes of lowering demand during a utility called Demand Response (DR) event. The testbed facility was the Foley Federal Building and US Courthouse ("Foley Federal Building") located in Las Vegas, NV. This is a 209,496 sq. ft. building constructed in the 1960s with major renovations in 2004. The facility was a good candidate due to the large prevalence of office and courthouse spaces in the GSA portfolio of buildings. It also has many features which allow integration into and control of the building and a strong facilities team to assist with the study. Quantitative and qualitative performance objectives were developed using GSA's GPG GEB project template along with input from the vendor and building facility staff; these are outlined in Table 1. The quantitative performance objectives focused on continuous demand management, energy efficiency, and automated demand response. The qualitative performance objectives focused on the ease of installation and commissioning as well as the operability of the GEB solution. Other performance metrics that are reported on include carbon reduction, cost effectiveness, and occupant acceptance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗