Search NASA⌕ Search

SEARCH · Search NASA

Results for “OpenAI Gym”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Hybrid-RL-MPC4CLR (Hybird-Reinforcement-Learning-Model-Predictive-Control-for-Reserve-Policy-Assisted-Critical-Load-Restoration-in-Distribution-Grids)

Hybrid-RL-MPC4CLR was developed as a hybrid controller for active distribution grid critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. The RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while the MPC models grid operations incorporating the RL policy actions (i.e., reserve requirements), renewable (wind and solar) power predictions, and load demand forecasts. The developers formulated the reserve requirement determination problem as a sequential decision-making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC simulation. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The software is developed using various software packages in Python. The MPC's optimal power flow (OPF) model is implemented using the Pyomo package, the RL simulation environment is implemented using the MPC simulation with various scenarios of renewable energy and load demand profiles and power outage beginning times, based on the OpenAI Gym framework. The RL agent training is performed using the RLlib Ray package. The RL algorithm is trained offline using historical forecasts of renewable generation and load demand profiles. Simulation analysis and performance tests are conducted using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine, and battery.

Eseye, Abinet Tesfaye↗

Reinforcement Learning for Intentional Islanding in Resilient Power Transmission Systems

Intentional islanding is the process of identifying and deliberately decomposing the transmission network to form self-sustained islands from an endangered network during disruptions to improve resilience and security. Most existing intentional islanding models are offline resilience decision tools and hence do not provide outage responses in a timely manner. In this paper, a reinforcement learning (RL) based model for intentional islanding is developed, which offers real-time switching control, online deployability, and adaptability to varying system conditions. The intentional islanding process is formulated as a Markov decision process, where the optimal transmission switching policy is learned using the RL approach. The control policy is learned over an environment that encompasses a Power System Simulator for Engineering (PSS/E) model of the transmission network, facilitated by an interface to the standard openAI Gym framework. The proposed RL-based methodology aims to form stable and self-sustainable islands by ensuring voltage stability while reducing the power mismatch in the formed islands. A proximal policy optimization algorithm is designed, which is suitable for controlling the on/off status of the switches with multi-layer perceptron as value and actor networks. The effectiveness of the proposed framework in the self-recovery of the grid by island formation is applied on the modified IEEE 39-bus test network and validated by dynamic simulations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

OCHRE™ Gymnasium (ochre_gym) [SWR-23-47]

OCHRE™ Gymnasium is a Python framework for conducting reinforcement learning research on residential building energy management problems. It is a lightweight wrapper around OCHRE, a Python-based thermal-electric residential building simulator developed in-house at NREL. OCHRE Gymnasium adheres to the standard OpenAI Gym API for a reinforcement learning environment. Detailed Sphinx documentation will be auto-generated and accompany this codebase. This documentation will detail the API, provide examples of how to use the code, etc.

Emami, Patrick↗

Deep reinforcement learning based optimization for a tightly coupled nuclear renewable integrated energy system

New ways to integrate energy systems to maximize efficiency are being sought to meet carbon emissions goals. Nuclear-renewable integrated energy system (NR-IES) concepts are a leading solution that couples a nuclear power plant with renewable energy, hydrogen generation plants, and energy storage systems, such that thermal and electrical power are dispatchable to fulfill grid-flexibility requirements while also producing hydrogen and maximizing revenue. Here, this paper introduces a deep reinforcement learning (DRL)-based framework to address the complex decision-making tasks for NR-IES. The objective is to maximize revenue by generating and selling hydrogen and electricity simultaneously according to their time-varying prices while keeping the energy flow in the subsystems in balance. A Python-based simulator for a NR-IES concept has been developed to integrate with OpenAI Gym and Ray/RLlib to enable an efficient and flexible computational framework for DRL research and development. Three state-of-the-art DRL algorithms have been investigated, including two-delayed deep deterministic policy gradient (TD3), soft-actor critic (SAC), proximal policy optimization (PPO), to illustrate DRL’s superiority for controlling NR-IES by comparing it with a conventional control approach, particle swarm optimization (PSO). In this effort, PPO has shown more-stable performance and also better generalization capability than SAC and TD3. Comparisons with PSO have demonstrated that, on average, PPO can achieve 13.9% more mean episode returns from the training process and 29.4% more mean episode returns from the testing process when different hydrogen-production targets are applied.

08 HYDROGEN↗

A high-fidelity building performance simulation test bed for the development and evaluation of advanced controls

We present an open-source building performance simulation test bed, the Advanced Controls Test Bed (ACTB), that interfaces high-fidelity Spawn of EnergyPlus building models, with advanced controllers implemented in Python. Additionally, the ACTB leverages the Building Optimization Testing and Alfalfa platforms for managing simulations, providing an external clock, a representational state transfer (REST) application programming interface (API), and key performance indicators for evaluating the effectiveness of control strategies. The REST API allows the development of external controllers programmed in languages such as Python, which provides flexibility and a rich choice of scientific libraries for designing control sequences. We present three test cases based on the U.S. Department of Energy's Reference Small Office Building to demonstrate the ACTB's capabilities: (a) rule-based controls compliant with ASHRAE Guideline 36 control sequences; (b) an economic model predictive control implemented using do-mpc; and (c) a deep Q-network reinforcement learning agent implemented using OpenAI Gym.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Enhancing Cyber Resilience of Networked Microgrids using Vertical Federated Reinforcement Learning

This paper presents a novel federated reinforcement learning (Fed-RL) methodology to inject sufficient resiliency into the operations of the network of microgrids. We consider adversarial actions to the voltage and power control loop reference signals at the grid forming (GFM) inverters in the microgrids which are essential to integrate renewable resources. Therefore, we formulate a resilient reinforcement learning training setup that uses these adversarial injections to generate episodic trajectories and train the RL agents to alleviate their impact on performance. To circumvent the concerns about data-sharing and privacy for different owners of the microgrids in the networked setting, we bring in the aspects of the federated operation to propose novel Fed-RL algorithms. As the dynamics of each microgrid are coupled due to electrical interlinks, the conventional federated RL approaches using decoupled independent environments are not applicable, which leads us to propose a multi-agent vertically federated variation of actor-critic algorithms, namely federated soft actor-critic (FedSAC). We have performed numerical simulations on an IEEE 123-bus benchmark test feeder with three microgrids by creating a customized simulation setup by encapsulating the microgrid dynamic simulations in GridLAB-D/HELICS co-simulation platform with the OpenAI Gym environment and validated the proposed resilient and secured learning methodology.

Artificial Intelligence (AI), reinforcement learni↗

Resilient Control of Networked Microgrids Using Vertical Federated Reinforcement Learning: Designs and Real-Time Test-Bed Validations

Improving system-level resiliency of networked microgrids against adversarial cyber-attacks is an important aspect in the current regime of increased inverter-based resources (IBRs). To achieve that, this paper contributes in designing a hierarchical control layer, in conjunction with the existing control layers, resilient to adversarial attack signals. Considering model complexities, unknown dynamical behaviors of IBRs, and privacy issues regarding data sharing in multi-party-owned microgrids, designing such a control layer is non-trivial. Here, to tackle these issues, a novel federated reinforcement learning (Fed-RL) method is proposed. To grasp the interconnected dynamics of networked microgrids, the paper develops Federated Soft Actor-Critic (FedSAC) algorithm following the vertical structure of implementing Fed-RL. Next, utilizing the OpenAI Gym interface, we built a custom set-up in GridLAB-D/HELICS co-simulation platform, named Resilient RL Co-simulation (ResRLCoSIM), to train the RL agents with IEEE 123-bus benchmark comprising 3 interconnected microgrids. Finally, the learned policies in the simulation are transferred to the real-time hardware-in-the-loop (HIL) test-bed developed using the high-fidelity Hypersim platform. Finally, experiments show that the simulator-trained RL controllers achieve desirable performance with the test-bed platform, validating the minimization of the sim-to-real gap.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Training reinforcement learning models via an adversarial evolutionary algorithm

When training for control problems, more episodes used in training usually leads to better generalizability, but more episodes also requires significantly more training time. There are a variety of approaches for selecting the way that training episodes are chosen, including fixed episodes, uniform sampling, and stochastic sampling, but they can all leave gaps in the training landscape. In this work, we describe an approach that leverages an adversarial evolutionary algorithm to identify the worst performing states for a given model. We then use information about these states in the next cycle of training; this process can be repeated until the desired level of model performance is met. We demonstrate this approach with the OpenAI Gym cart-pole problem. With this problem, we show that the adversarial evolutionary algorithm did not reduce the number of episodes required in training needed to attain model generalizability when compared with stochastic sampling, and actually performed slightly worse.

Coletti, Mark↗

Graph-Env

The Graph-Env library provides a framework for adapting graph search problems into OpenAI Gym environments for reinforcement learning. In other words, Graph-Env enables reinforcement learning algorithms to be applied to graph search problems. Graph search problems include molecule and crystal structure design problems; route planning problems, including traveling salesperson problem; puzzles and games, such as chess and go; shortest path problems; minimum spanning tree; vehicle path generation; and others.

Tripp, Charles↗

RLC4CLR (Reinforcement Learning Controller for Critical Load Restoration Problems)

RLC4CLR demonstrates using a reinforcement learning controller (RLC) to solve a critical load restoration (CLR) problem, which improves the grid resilience after a substation outage event. RLC4CLR consists of two parts. (1) RL environment: This environment encapsulates the CLR problem to be solved and provides interfacing functions to follow the standard OpenAI Gym format. A power system simulator, i.e., OpenDSS, is included to provide the power flow solution. Controller inputs and outputs (RL state and action) as well as the reward are defined in this environment as well. In summary, the RL environment is the problem formulation from which the RL agent can learn. (2) RL training script: The training script enables the RL agent to learn its control policy by interacting with the RL environment. For RL training, an open-sourced RL library, i.e., RLlib, is leveraged which is based on a distributed computing framework (Ray). The training script is designed to be able to be run on both local machine or the NREL HPC system. Other components of RLC4CLR include input data, e.g., grid model (standard IEEE test feeders), and other files used for results analysis.

Zhang, Xiangyu↗

A Hybrid Reinforcement Learning-MPC Approach for Distribution System Critical Load Restoration: Preprint

This paper proposes a hybrid control approach for distribution system critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while MPC models grid operations incorporating RL policy actions, i.e., the reserve requirement, renewable (wind and solar) power predictions, and load demand forecasts. We formulate the reserve requirement determination problem as a sequential decision making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The RL algorithm is trained off-line using historical forecast of renewable generation and load demand. The method is tested using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine and battery. Case studies demonstrated that the proposed method outperforms other operating reserve determination methods.

distribution system↗

A Hybrid Reinforcement Learning-MPC Approach for Distribution System Critical Load Restoration

This paper proposes a hybrid control approach for distribution system critical load restoration, combining deep reinforcement learning (RL) and model predictive control (MPC) aiming at maximizing total restored load following an extreme event. RL determines a policy for quantifying operating reserve requirements, thereby hedging against uncertainty, while MPC models grid operations incorporating RL policy actions (i.e., reserve requirements), renewable (wind and solar) power predictions, and load demand forecasts. We formulate the reserve requirement determination problem as a sequential decision-making problem based on the Markov Decision Process (MDP) and design an RL learning environment based on the OpenAI Gym framework and MPC simulation. The RL agent reward and MPC objective function aim to maximize and monotonically increase total restored load and minimize load shedding and renewable power curtailment. The RL algorithm is trained offline using a historical forecast of renewable generation and load demand. The method is tested using a modified IEEE 13-bus distribution test feeder containing wind turbine, photovoltaic, microturbine, and battery. Case studies demonstrated that the proposed method outperforms other policies with static operating reserves.

distribution system↗

Ten questions concerning reinforcement learning for building energy management

As buildings account for approximately 40% of global energy consumption and associated greenhouse gas emissions, their role in decarbonizing the power grid is crucial. The increased integration of variable energy sources, such as renewables, introduces uncertainties and unprecedented flexibilities, necessitating buildings to adapt their energy demand to enhance grid resiliency. Consequently, buildings must transition from passive energy consumers to active grid assets, providing demand flexibility and energy elasticity while maintaining occupant comfort and health. This fundamental shift demands advanced optimal control methods to manage escalating energy demand and avert power outages. Reinforcement learning (RL) emerges as a promising method to address these challenges. Here, in this paper, we explore ten questions related to the application of RL in buildings, specifically targeting flexible energy management. We consider the growing availability of data, advancements in machine learning algorithms, open-source tools, and the practical deployment aspects associated with software and hardware requirements. Our objective is to deliver a comprehensive introduction to RL, present an overview of existing research and accomplishments, underscore the challenges and opportunities, and propose potential future research directions to expedite the adoption of RL for building energy management.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing

Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. Furthermore, the experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.

97 MATHEMATICS AND COMPUTING↗

PowerGridworld: A Framework for Multi-Agent Reinforcement Learning in Power Systems: Preprint

We present the PowerGridworld software package to provide users with a light-weight, modular, and customizable framework for creating power systems-focused, multi-agent gym environments that readily integrate with existing training frameworks for reinforcement learning (RL). While many frameworks exist for training multi-agent (MA) RL policies, none exist to rapidly prototype and develop the environments themselves, especially in the context of heterogeneous (composite, multi-device) power systems where power flow solutions are required to define grid-level variables and costs. PowerGridworld is an open-source software package that helps to fill this gap. To highlight PowerGridworld's key features, we present two case studies and demonstrate learning multi-agent RL policies using both OpenAI's MADDPG and RLLib's PPO algorithms where, in both cases, at least some subset of agents incorporate elements of the power flow solution at each time step as part of their reward (negative cost) structures.

MATHEMATICS AND COMPUTING↗