Search NASA⌕ Search

SEARCH · Search NASA

Results for “on-policy learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

On-policy learning-based deep reinforcement learning assessment for building control efficiency and stability

Artificial intelligence technologies have emerged as a game changer not only in specific applications such as image recognition and machine translation but also in many scientific domains. In particular, as deep reinforcement learning (DRL) has shown great success in complex control problems, DRL-based control has been considered as a potential solution to efficiently control and manage building systems. However, broad assessment of DRL-based building control is still required to characterize their pros and cons in comparison with conventional building control methods (e.g., rule-based feedback controls). In this paper, we assessed DRL-based controls with on-policy learning-based algorithms and continuous control actions for cooling control of large office buildings in the summer season to minimize whole-building energy use and occupant discomfort. We compared DRL-based control methods with two baseline control methods: (1) a pre-determined schedule with supply temperature and static pressure setpoints, and (2) advanced reset method that adjusts setpoints based on heuristic rules, i.e., ASHRAE Guideline 36. We also tested the DRL algorithms to evaluate their performances in multiple climate locations. We found that DRL-based control methods outperformed the baseline control methods in terms of energy savings while maintaining a thermal comfort. DRL reduced energy use between ~4%–22% on average compared to the baseline methods, depending on climate location. We also evaluated DRL-based control in terms of control stability and showed that DRL-based methods should address the span of hardware lifetimes in practical operations.

control stability↗

PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control

This paper develops an efficient multi-agent deep reinforcement learning algorithm for cooperative controls in powergrids. Specifically, we consider the decentralized inverter-based secondary voltage control problem in distributed generators (DGs), which is first formulated as a cooperative multi-agent reinforcement learning (MARL) problem. We then propose a novel on-policy MARL algorithm, PowerNet, in which each agent (DG) learns a control policy based on (sub-)global reward but local states and encoded communication messages from its neighbors. Motivated by the fact that a local control from one agent has limited impact on agents distant from it, we exploit a novel spatial discount factor to reduce the effect from remote agents, to expedite the training process and improve scalability. Furthermore, a differentiable, learning-based communication protocol is employed to foster the collaborations among neighboring agents. In addition, to mitigate the effects of system uncertainty and random noise introduced during on-policy learning, we utilize an action smoothing factor to stabilize the policy execution. To facilitate training and evaluation, we develop PGSim, an efficient, high-fidelity powergrid simulation platform. Here, experimental results in two microgrid setups show that the developed PowerNet outperforms the conventional model-based control method, as well as several state-of-the-art MARL algorithms. The decentralized learning scheme and high sample efficiency also make it viable to large-scale power grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Sequential Decision Making (SDM) for Mesh Refinement and Model Selection in Multiscale, Multi-Physics Applications

Intelligent automation and decision support are needed to enhance computational efficiency and robustness in multiscale and multi-physics problems, including materials science, manufacturing, and climate and weather modeling. Current scientific computing approaches for enabling decisions by scientists fail to explore the role of learning, reasoning, and probabilistic planning. Often these decisions are not performed in real-time during the computation but are made prior to the start of the computation, which must be interrupted in order to make changes to the prior choices. Such interruptions at different stages of the computation increase the total computing time and the need for a human expert to frequently monitor the results. State of art scientific computing methods consist of rule-based algorithms that cannot automatically adapt to a dynamically changing computing environment. The development of a Sequential Decision Making (SDM) framework will automate scientific computing by optimizing the policies for mesh refinement, time-stepping, model and algorithm selection, resource allocation, and pre and post-processing. Our agent SDM framework for scientific computing will consist of data-driven learning (Classifier), automated reasoning (contextual knowledge), and probabilistic planning (Reinforcement Learning). In this project, we focused on three problems to demonstrate our SDM framework on a set of ordinary and partial differential equations. Classification of Lorenz system regions using Feed-Forward Neural Networks examined learning in the SDM framework. On the other hand, reasoning and planning in the SDM framework were used in two problems: adaptive time-stepping for nonlinear ODEs using on-policy RL algorithms, and adaptive mesh refinement for 2-D PDEs using off-policy RL algorithms.

97 MATHEMATICS AND COMPUTING↗

Optimizing Patient-Specific Medication Regimen Policies Using Wearable Sensors in Parkinson’s Disease

Effective treatment of Parkinson’s disease (PD) is a continual challenge for healthcare providers, and providers can benefit from leveraging emerging technologies to supplement traditional clinic care. We develop a data-driven reinforcement learning (RL) framework to optimize PD medication regimens through wearable sensors. We leverage a data set of n = 26 PD patients who wore wrist-mounted movement trackers for two separate six-day periods. Using these data, we first build and validate a simulation model of how individual patients’ movement symptoms respond to medication administration. We then pair this simulation model with an on-policy RL algorithm that recommends optimal medication types, timing, and dosages during the day while incorporating human-in-the-loop considerations on medication administration. The results show that the RL-prescribed medication regimens outperform physicians’ medication regimens, despite physicians having access to the same data as the RL agent. To validate our results, we assess our wearable-based RL medication regimens using n = 399 PD patients from the Parkinson’s Progression Markers Initiative data set. We show that the wearable-based RL medication regimens would lead to significant symptom improvement for these patients, even more so than training RL policies directly from this data set. In doing so, we show that RL models from even small data sets of wearable data can offer novel, generalizable clinical insights and medication strategies, which may outperform those derived from larger data sets without wearable data.

60 APPLIED LIFE SCIENCES↗

Approximate Dynamic Programming With Enhanced Off-Policy Learning for Coordinating Distributed Energy Resources

Herein this paper proposes an innovative approximate dynamic programming (ADP) method for distributed energy resource coordination with the loss of life of battery energy storage system (BESS) explicitly modeled. The dispatch policy is designed to account for both calendrical and cyclical aging effects on BESS, explicitly modeling the impacts of ambient temperature on BESS lifespan. The proposed ADP employs an adaptive critic method and enhanced off-policy deterministic policy gradient (DPG) strategy, addressing the limitations of the on-policy gradient-based ADP approaches, including inadequate exploration, low data usage, and computational complexity. In particular, a customized policy is proposed to guide the algorithm to explore some promising decisions and thereby improve exploration capability and learning efficiency compared to conventional DPG-based learning approaches, which may struggle to find a global optimum due to random noisy action-based exploration or require expert demonstration with extra effort. The proposed method is illustrated using the IEEE 123-node system and compared with the existing ADP methods to prove solution accuracy and demonstrate the effects of incorporating degradation models into control design. Case studies showed that the proposed ADP effectively coordinates DERs with a 10 times smaller optimization gap compared to existing methods, and the incorporation of the BESS life loss model ensures the expected lifespan and results in significant cost savings.

24 POWER TRANSMISSION AND DISTRIBUTION↗