Search NASA⌕ Search

SEARCH · Search NASA

Results for “multiagent reinforcement learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Interpreting Primal-Dual Algorithms for Constrained Multiagent Reinforcement Learning

Constrained multiagent reinforcement learning (C-MARL) is gaining importance as MARL algorithms find new applications in real-world systems ranging from energy systems to drone swarms. Most C-MARL algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of this penalty term on the MARL problem. First, we show that the standard practice of using the constraint function as the penalty leads to a weak notion of safety. However, by making simple modifications to the penalty term, we can enforce meaningful probabilistic (chance and conditional value at risk) constraints. Second, we quantify the effect of the penalty term on the value function, uncovering an improved value estimation procedure. We use these insights to propose a constrained multiagent advantage actor critic (C-MAA2C) algorithm. Simulations in a simple constrained multiagent environment affirm that our reinterpretation of the primal-dual method in terms of probabilistic constraints is effective, and that our proposed value estimate accelerates convergence to a safe joint policy.

chance constraints↗

Interpreting Primal-Dual Algorithms for Constrained Multiagent Reinforcement Learning: Preprint

We study multiagent reinforcement learning (MARL) with constraints. This setting is gaining importance as MARL algorithms find new applications in real-world systems ranging from power grids to drone swarms. Most constrained MARL (C-MARL) algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of the primal-dual approach on the constraints and value function. First, we show that using the constraint evaluation as the penalty leads to a weak notion of safety, but by making simple modifications to the penalty function, we can enforce meaningful probabilistic safety constraints. Second, we show that the penalty term changes the value function in a way that is easy to model, and demonstrate the consequences of not doing so. We conclude with simulations in a simple constrained multiagent environment to back up the theoretical results.

data-driven control↗

Large-Eddy Simulation of Flow Over Boeing Gaussian Bump Using Multiagent Reinforcement Learning Wall Model: Preprint

We develop a wall model for large-eddy simulation (LES) that takes into account various pressure-gradient effects using multi-agent reinforcement learning. The model is trained using low-Reynolds-number flow over periodic hills with agents distributed on the wall at various computational grid points. It utilizes a wall eddy-viscosity formulation as the boundary condition to apply the modeled wall shear stress. Each agent receives states based on local instantaneous flow quantities at an off-wall location, computes a reward based on the estimated wall-shear stress, and provides an action to update the wall eddy viscosity at each time step. The trained wall model is validated in wall-modeled LES of flow over periodic hills at higher Reynolds numbers, and the results show the effectiveness of the model on flow with pressure gradients. The analysis of the trained model indicates that the model is capable of distinguishing between the various pressure gradient regimes present in the flow. To further assess the robustness of the developed wall model, simulations of flow over the Boeing Gaussian bump are conducted at a Reynolds number of 2 x 10^6, based on the free-stream velocity and the bump width. The results of mean skin friction and pressure on the bump surface, as well as the velocity statistics of the flow field, are compared to those obtained from equilibrium wall model (EQWM) simulations and published experimental data sets. The developed wall model is found to successfully capture the acceleration and deceleration of the turbulent boundary layer on the bump surface, providing better predictions of skin friction near the bump peak and exhibiting comparable performance to the EQWM with respect to the wall pressure and velocity field. We also conclude that the subgrid-scale model is crucial to the accurate prediction of the flow field, in particular the prediction of separation.

boundary layer↗

Multi-Agent Reinforcement Learning for Distribution System Critical Load Restoration

Grid resilience has become a critical topic recently because of the increasing occurrence of extreme events and the growing integration of intermittent renewable energy sources. To build a resilient distribution system, this paper develops a multiagent reinforcement learning-based (MARL) method to coordinate distribution energy resources (DERs) dispatch, load pickup, and network reconfiguration for load restoration after a system outage. With the help of two types of control agents, namely critical load restoration (CLR) and coordination (COR) agents, system loads can be restored efficiently, given available resources. The effectiveness and superiority of the proposed algorithm are demonstrated through simulations and comparative studies on a real distribution feeder in Western Colorado.

distribution system↗

A Novel LDPP-MADDPG Approach for Distributed Power Allocation in mmWave Cellular Networks

This paper considers the problem of distributed beam scheduling and power allocation problem in millimeter- Wave (mmWave) cellular networks, in which multiple Base Stations (BSs) operate as individual operators over a shared spectrum. We propose a novel learning-aided approach that integrates the Lyapunov Drift-Plus-Penalty (LDPP) framework and Multi-agent Deep Deterministic Policy Gradient (MADDPG) reinforcement learning algorithms. This offers a powerful approach to learning stable and constraint-aware policies, reaping the joint benefit of both LDPP and MADDPG, in complex multiagent environments. The major challenge for this approach is to integrate these two approaches in a meaningful and effective manner. The key idea to solve this problem is to introduce a novel feature of local observation that incorporates potential negative value of the reward function due to the stochastic constraints introduced by the LDPP framework. Empirical results demonstrate that our proposed scheme outperforms the baseline methods under various conditions.

99 - GENERAL AND MISCELLANEOUS↗