Search NASA⌕ Search

Engineering topics

Chen, B

Publications and source records attributed to Chen, B.

AlphaBuilding ResCommunity: A multi-agent virtual testbed for community-level load coordination

Training and validating algorithms in a simulation testbed can accelerate research and applications of optimal control of residential loads to improve energy flexibility and grid resilience. We developed an open-source simulation environment, AlphaBuilding ResCommunity, that can be used to train and validate algorithms to control a single thermostatically controlled load (TCL) or coordinate a group of TCLs. We used reduced-order models to simulate the thermodynamics of TCLs, and the parameter values were determined from the connected smart thermostat data of real households. The environment was built upon the standardized OpenAI Gym interface. Ancillary functions, such as retrieving the parameters and weather forecasts, are provided to facilitate control strategies that require predictive information. Compared with existing efforts, AlphaBuilding ResCommunity has three advantages: (1) more realistic model settings because the parameter values are identified from actual household operating data, and modelling and measurement uncertainty are considered; (2) passive thermal storage control; and (3) ease of use due to a simple software dependency and standardized interface. We demonstrated the applications of the environment by implementing a Kalman Filter and Model Predictive Control on a single TCL and a Priority-Stack-Based Control and Alternating Direction Method of Multipliers to coordinate multiple TCLs for load tracking.

Wang, Z↗

Towards Off-policy Evaluation as a Prerequisite for Real-world Reinforcement Learning in Building Control

We present an initial study of off-policy evaluation (OPE), a problem prerequisite to real-world reinforcement learning (RL), in the context of building control. OPE is the problem of estimating a policy's performance without running it on the actual system, using historical data from the existing controller. It enables the control engineers to ensure a new, pretrained policy satisfies the performance requirements and safety constraints of a real-world system, prior to interacting with it. While many methods have been developed for OPE, no study has evaluated which ones are suitable for building operational data, which are generated by deterministic policies and have limited coverage of the state-action space. After reviewing existing works and their assumptions, we adopted the approximate model (AM) method. Furthermore, we used bootstrapping to quantify uncertainty and correct for bias. In a simulation study, we evaluated the proposed approach on 10 policies pretrained with imitation learning. On average, the AM method estimated the energy and comfort costs with 1.84% and 14.1% error, respectively.

Chen, B↗