Search NASA⌕ Search

SEARCH · Search NASA

Results for “deep Q-networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Deep Reinforcement Learning for Autonomous Water Heater Control

Electric water heaters represent 14% of the electricity consumption in residential buildings. An average household in the United States (U.S.) spends about USD 400–600 (0.45 ¢/L–0.68 ¢/L) on water heating every year. In this context, water heaters are often considered as a valuable asset for Demand Response (DR) and building energy management system (BEMS) applications. To this end, this study proposes a model-free deep reinforcement learning (RL) approach that aims to minimize the electricity cost of a water heater under a time-of-use (TOU) electricity pricing policy by only using standard DR commands. In this approach, a set of RL agents, with different look ahead periods, were trained using the deep Q-networks (DQN) algorithm and their performance was tested on an unseen pair of price and hot water usage profiles. The testing results showed that the RL agents can help save electricity cost in the range of 19% to 35% compared to the baseline operation without causing any discomfort to end users. Additionally, the RL agents outperformed rule-based and model predictive control (MPC)-based controllers and achieved comparable performance to optimization-based control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deep Reinforcement Learning based Model-free On-line Dynamic Multi-Microgrid Formation to Enhance Resilience

Multi-microgrid formation (MMGF) is a promising solution for enhancing power system resilience. This paper proposes a new deep reinforcement learning (RL) based model-free on-line dynamic MMGF scheme. Additionally, the dynamic MMGF problem is formulated as a Markov decision process, and a complete deep RL framework is specially designed for the topologytransformable micro-grids. In order to reduce the large action space caused by flexible switch operations, a topology transformation method is proposed and an action-decoupling Q-value is applied. Then, a convolutional neural network (CNN) based multi-buffer double deep Q-network (CM-DDQN) is developed to further improve the learning ability of the original DQN method. The proposed deep RL method provides real-time computing to support the on-line dynamic MMGF scheme, and the scheme handles a long-term resilience enhancement problem using an adaptive on-line MMGF to defend changeable conditions. The effectiveness of the proposed method is validated using a 7-bus system and the IEEE 123-bus system. The results show strong learning ability, timely response for varying system conditions and convincing resilience enhancement.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Reinforcement Learning for Real-Time Hardware-Based Energy System Experiments: Preprint

In the context of urgent climate challenges and the pressing need for rapid technology development, Reinforcement Learning (RL) stands as a compelling data-driven method for controlling real-world physical systems. However, RL implementation often entails time-consuming and computationally intensive data collection and training processes, rendering them inefficient for real-time applications that lack non-real-time models. To address these limitations, real-time emulation techniques have emerged as valuable tools for the lab-scale rapid prototyping of intricate energy systems. While emulated systems offer a bridge between simulation and reality, they too face constraints, hindering comprehensive characterization, testing, and development. In this research, we construct a surrogate model using limited data from simulated systems, enabling an efficient and effective training process for a Double Deep Q-Network (DDQN) agent for future deployment. Our approach is illustrated through a hydropower application, demonstrating the practical impact of our approach on climate-related technology development.

deep Q-learning↗

A Cyber-Physical System for Freeway Ramp Meter Signal Control Using Deep Reinforcement Learning in a Connected Environment

Freeway bottlenecks such as on-ramp merging areas account for about 40% of recurring freeway congestion. It is generally agreed that building more roads and adding more lanes to existing infrastructure does not solve the congestion problem, and so dynamic traffic control measures offer a more cost-effective alternative. Ramp meters, traffic signal devices that regulate traffic flow entering freeways, are among the most effective measures to mitigate congestion at on-ramp merging areas on freeways. The confluence of deep reinforcement learning (RL) and connectivity provides a possible solution to advance ramp meter signal control. Deep RL is a group of machine-learning methods that enables an agent learning from the environment to improve its performance. In this study, three deep RL methods-proximal policy optimization (PPO), Ape-X deep Q-network (DQN), and asynchronous advantage actor-critic agents (A3C)-are explored for ramp meter signal control to maximize vehicle speed and traffic throughput, as well as to minimize energy consumption and emissions at freeway on-ramp merging areas in a connected environment. The low computational requirement and scalability of deep RL for deployment make it a powerful optimization tool for time-sensitive applications such as ramp meter signal control. The results of this study show that deep RL methods yield superior performance to both a fixed-time controller and ALINE A, a state-of-the-art feedback controller.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

Optimal Coordination of Electric Vehicles for Grid Services using Deep Reinforcement Learning

Recent research has shown the effectiveness of reinforcement learning (RL) in coordinating electric vehicles (EVs) with vehicle-to-grid capabilities for grid services. However, many of these studies rely on lookup table and deep Q-network techniques, which can be impractical when dealing with continuous states and actions. In addition, existing RL designs inadequately account for battery aging effects, EV user satisfaction, uncertain departure and arrival time, and trip distance, which may compromise effective coordination. This paper aims to bridge these gaps by developing an innovative deep deterministic policy gradient-based RL framework for optimal coordination of EVs. Case studies were carried out using a test system with 100 EVs, and numerical analysis results showed that the proposed RL framework can effectively coordinate EVs to maximize economic benefits and user satisfaction while ensuring the expected battery lifespan.

Das, Avijit↗

Designing reinforcement learning algorithms for building HVAC control: From experimental observation to simulation comparisons

Advanced supervisory-level control with reinforcement learning (RL) is regarded as a promising solution for HVAC systems to minimize energy consumption while maintaining thermal comfort and indoor air quality. However, most RL applications were conducted in the simulation environment rather than real-world HVAC systems. This paper developed a value-based RL controller termed Deep Q-Network (DQN) for a typical central HVAC system and evaluated its performance in a building test facility. By comparing DQN with a rule-based controller, the study not only demonstrated the cases where DQN could properly maintain indoor comfort but also discussed possible reasons why DQN failed in some other situations. Recognizing the limitations of value-based RL algorithms from the experimental tests, a simulation study was conducted to compare DQN with an alternative RL approach, an actor–critic algorithm termed Deep Deterministic Policy Gradient (DDPG). In scenarios with a relatively large action space, DDPG outperformed DQN by requiring fewer computational resources and achieving better thermal comfort, lower energy consumption, and more stable control actions. The findings suggest that the ability of DDPG to handle continuous control variables more effectively allows for faster convergence in training and more precise control in practice, which enhances the overall efficiency and reliability of the HVAC system.

Guo, Fangzhou↗

Deep Reinforcement Learning based Routing in an Air-to-Air Ad-hoc Network

This paper studies the Multiple Sources and Multiple Destinations (MSMD) routing problem in a dynamic Air-to-Air Ad-hoc Network (AAAN). We consider a spectrum limited scenario where multiple links have to share the same frequency channel so that co-channel interference becomes inevitable. As a result, routing decisions and spectrum access are coupled and must be jointly considered. This paper proposes a deep Q-learning based algorithm to find an optimal routing and channel selection strategy that minimizes the end-to-end communication delay. Specifically, under the assumption that only local information is available to every node, the Deep Q-Network (DQN) is trained offline to learn the optimal routing and channel selection strategy. After the trained DQN is implemented in every node, multiple relay nodes can simultaneously determine their next-hop relay and channel selections in real-time. Simulation results demonstrate the efficacy of our proposed algorithm.

AAAN↗

Deep Reinforcement Learning based Routing in an Air-to-Air Ad-hoc Network

This paper studies the Multiple Sources and Multiple Destinations (MSMD) routing problem in a dynamic Air-to-Air Ad-hoc Network (AAAN). We consider a spectrum limited scenario where multiple links have to share the same frequency channel so that co-channel interference becomes inevitable. As a result, routing decisions and spectrum access are coupled and must be jointly considered. This paper proposes a deep Q-learning based algorithm to find an optimal routing and channel selection strategy that minimizes the end-to-end communication delay. Specifically, under the assumption that only local information is available to every node, the Deep Q-Network (DQN) is trained offline to learn the optimal routing and channel selection strategy. After the trained DQN is implemented in every node, multiple relay nodes can simultaneously determine their next-hop relay and channel selections in real-time. Simulation results demonstrate the efficacy of our proposed algorithm.

AAAN↗

Deep Reinforcement Learning Based Smart Water Heater Control for Reducing Electricity Consumption and Carbon Emission

Water heating is the third largest electricity consumer in U.S. households, after space heating and cooling. Thus, water heaters represent a significant potential for reducing electricity consumption and associated CO2 emissions of residential buildings. To this end, this study proposes a model-free deep reinforcement learning (RL) approach that aims to minimize the electricity consumption and the CO2 emissions of a heat pump water heater without affecting user comfort. In this approach, a set of RL agents focusing on either electricity saving or emission reduction, with different look ahead periods, were trained using the deep Q-networks (DQN) algorithm and their performance was tested on different hot water usage and Marginal Operating Emissions Rate (MOER) profiles. The testing results showed that the RL agents that focus on electricity saving can save electricity in the range of 12–22% by operating the water heater with maximum heat pump efficiency and minimum electric element utilization. On the other hand, the RL agents that focus on emission reduction reduced emissions in the range of 18–37% by making use of the variable MOER values. These RL agents used the heat pump and/or an element when the MOER values are low due to the availability of renewable energy sources (e.g., solar and wind) and mostly avoided the periods of carbon-intensive periods. Overall, these results showed that the proposed RL approach can help minimize the electricity consumption and the CO2 emissions of a heat pump water heater without having any prior knowledge about the device.

Amasyali, Kadir↗

Deep reinforcement learning assisted co-optimization of Volt-VAR grid service in distribution networks

With the increasing penetration of distributed energy resources in distribution networks, Volt-VAR control and optimization (VVC/VVO) have become very important to ensure an acceptable quality of service to all customers. System operators can rely on slow-responding utility devices, including capacitor banks and on-load tap changing transformers, along with fast-responding battery and photovoltaic (PV) inverters for the VVC/VVO implementation. Because of variations in response time of these two classes of devices, and different control actions (discrete versus continuous), coordinated and optimal scheduling and operation have become of utmost importance. Here, this paper develops a look-ahead deep reinforcement learning (DRL)-based multi-objective VVO technique to improve the voltage profile of active distribution networks, decrease network and inverter power loss, and save the operational cost of the grid. It proposes a deep deterministic policy gradient (DDPG)-based approach to schedule the optimal reactive and/or active power set-points of fast-responding inverters, and a deep Q-network (DQN)-based DRL agent to schedule the discrete decisions variables of slow-responding assets. The reactive power output of PV and battery smart inverters are scheduled at 30-minute intervals and the capacitors’ commitment status is scheduled with several hour intervals. The proposed framework is validated on the modified IEEE 34-bus and 123-bus test cases with embedded PV and PV-plus-storage. To validate the efficacy of the proposed VVO, it is compared with several scenarios, including the base case without VVO, localized droop control of DERs, DDPG-only, and twin delayed DDPG (TD3) agent-based DRL techniques. The results justify the superior performance of the proposed method to improve the voltage profile, reduce network power loss, and minimize the look-ahead grid operational cost while minimizing the undesirable power losses in inverters as a result of power factor adjustments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decentralised Reinforcement Learning for Dynamic Cyberattack Response in Microgrid Networks

Microgrids rely on communication networks for reliable operation, which makes them inherently vulnerable to cyberattacks. Such attacks can destabilise system dynamics and drive states away from their nominal operating trajectories. Although several physics-informed and machine learning-based strategies have been developed to counter these threats, the rapidly evolving cyber landscape enables adversaries to bypass static defences or rules-based mitigation approaches. This paper proposes a dynamic, online-trained and fully decentralised reinforcement learning (RL)-based cyberattack response framework to protect microgrids from evolving cyberattacks. The proposed framework deploys multiple deep Q-networks (DQNs), each associated with a distributed energy resource (DER), to enable localised and adaptive attack mitigation. In this framework, each DQN processes local voltage and frequency measurements—combined with intrusion detection system (IDS) alerts—as observations and rewards to guide decision-making. Extensive simulation studies demonstrate the robustness of the proposed framework under diverse attack scenarios and varying IDS-induced detection delays. Comparative analysis highlights its superiority over existing static or preexisting rules-based mitigation approaches. Finally, we present an analysis that shows the framework's scalability to real-life microgrids with more interacting agents.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A high-fidelity building performance simulation test bed for the development and evaluation of advanced controls

We present an open-source building performance simulation test bed, the Advanced Controls Test Bed (ACTB), that interfaces high-fidelity Spawn of EnergyPlus building models, with advanced controllers implemented in Python. Additionally, the ACTB leverages the Building Optimization Testing and Alfalfa platforms for managing simulations, providing an external clock, a representational state transfer (REST) application programming interface (API), and key performance indicators for evaluating the effectiveness of control strategies. The REST API allows the development of external controllers programmed in languages such as Python, which provides flexibility and a rich choice of scientific libraries for designing control sequences. We present three test cases based on the U.S. Department of Energy's Reference Small Office Building to demonstrate the ACTB's capabilities: (a) rule-based controls compliant with ASHRAE Guideline 36 control sequences; (b) an economic model predictive control implemented using do-mpc; and (c) a deep Q-network reinforcement learning agent implemented using OpenAI Gym.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Integrated Routing and Traffic Signal Control for CAVs via Reinforcement Learning Approach

Incorporating Connected and Automated Vehicles (CAVs) into urban traffic networks presents opportunities and challenges for traffic management systems. This paper aims to develop an integrated routing and traffic signal control system designed explicitly for CAVs, utilizing a Reinforcement Learning (RL) approach. The objective is to enhance traffic flow and improve overall transportation efficiency in the controlled areas. We propose an innovative framework that employs the Deep Reinforcement Learning (DRL) algorithm, especially the Deep Q-network (DQN), to dynamically adjust the number of vehicles in the routes and the duration of traffic signals. Our simulation results demonstrate that a DQN agent successfully optimizes the number of vehicles in the routes and traffic signal timings of traffic signal controllers, eventually reducing total travel time. The study illustrates the potential usage of RL-based systems in managing routing and traffic signals for CAVs, offering a promising opportunity for future urban traffic management strategies.

Park, Jiho [New York University]↗

Buyers Collusion in Incentivized Forwarding Networks: A Multi-Agent Reinforcement Learning Study

We present the issue of monetarily incentivized forwarding in a multi-hop mesh network architecture from an economic perspective. It is anticipated that credit-incentivized forwarding and relaying will be a simple method of exchanging transmission power and spectrum for connectivity. However, gateways and forwarding nodes, like any other free market, may create an oligopolistic market for the users they serve. In this study, a coalition scheme between buyers aims to address price control by gateways or nodes closer to gateways. In a Stackelberg competition game, buyer agents (users) and sellers (gateways) make decisions using reinforcement learning (RL), with decentralized Deep Q-Networks to buy and sell forwarding resources. We allow communication links between the buyers with a limited messaging space, without defining a collusion mechanism. The idea is to demonstrate that through messaging, and RL tacit collusion can emerge between agents in a decentralized setup. The multi-agent reinforcement learning (MARL) system is presented and analyzed from a machine-learning perspective. Moreover, MARL dynamics are discussed via mean field analysis to better understand divergence causes and make implementation recommendations for such systems. Finally, the simulation results show the results of coordination among the users.

42 ENGINEERING↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Microservice Architecture for Cognitive Networks

This develops the concept of a cognitive network and describes a microservice based architecture which could be used to implement such a system. Delay tolerant networking (DTN) influences the design of the architecture as well as the networking scenarios that the system attempts to address. A cognitive storage and fragmentation service is developed based on existing artificial intelligence techniques such as Advantage Actor Critic (A2C) and Deep Q-Networks. The system is simulated using OpenAI Gym in a custom developed DTN environment.

cognitive networks↗

Self-Driving Telescopes: Autonomous Scheduling of Astronomical Observation Campaigns with Offline Reinforcement Learning

Modern astronomical experiments are designed to achieve multiple scientific goals, from studies of galaxy evolution to cosmic acceleration. These goals require data of many different classes of night-sky objects, each of which has a particular set of observational needs. These observational needs are typically in strong competition with one another. This poses a challenging multi-objective optimization problem that remains unsolved. The effectiveness of Reinforcement Learning (RL) as a valuable paradigm for training autonomous systems has been well-demonstrated, and it may provide the basis for self-driving telescopes capable of optimizing the scheduling for astronomy campaigns. Simulated datasets containing examples of interactions between a telescope and a discrete set of sky locations on the celestial sphere can be used to train an RL model to sequentially gather data from these several locations to maximize a cumulative reward as a measure of the quality of the data gathered. We use simulated data to test and compare multiple implementations of a Deep Q-Network (DQN) for the task of optimizing the schedule of observations from the Stone Edge Observatory (SEO). We combine multiple improvements on the DQN and adjustments to the dataset, showing that DQNs can achieve an average reward of 87%+-6% of the maximum achievable reward in each state on the test set. This is the first comparison of offline RL algorithms for a particular astronomical challenge and the first open-source framework for performing such a comparison and assessment task.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗