Search NASA⌕ Search

DOE OSTI · 2319191

Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning

Abstract

In multi-timescale multi-agent reinforcement learning (MARL), agents interact across different timescales. In general, policies for time-dependent behaviors, such as those induced by multiple timescales, are non-stationary. Learning non-stationary policies is challenging and typically requires sophisticated or inefficient algorithms. Motivated by the prevalence of this control problem in real-world complex systems, we introduce a simple framework for learning non-stationary policies for multi-timescale MARL. Our approach uses available information about agent timescales to define and learn periodic multi-agent policies. In detail, we theoretically demonstrate that the effects of non-stationarity introduced by multiple timescales can be learned by a periodic multi-agent policy. To learn such policies, we propose a policy gradient algorithm that parameterizes the actor and critic with phase-functioned neural networks, which provide an inductive bias for periodicity. The framework's ability to effectively learn multi-timescale policies is validated on a gridworld and building energy management environment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Emami, Patrick, Zhang, Xiangyu, Biagioni, David, Zamzam, Ahmed S.. 2024-01-19. Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning. https://doi.org/10.1109/cdc49753.2023.10384223

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Systematic multi-machine analysis of the exhaust time-dependent behavior in tokamaks

The understanding of the time-scales and associated transient behavior of fusion exhaust plasmas plays a crucial role in its dynamic modeling and its control. This work presents an overview of experimental investigations of the exhaust dynamics in TCV, MAST-U, ASDEX-Upgrade, WEST, DIII-D, and JET. From the presented experiments, a clear picture arises on properties of the exhaust dynamics across machines. Particularly, we observe that the scrape-off layer equilibrates on fast time-scales ($>$ 70 Hz) and that exhaust dynamics measured in response to gas valve modulations mostly behave smoothly and linearly, with similarities across devices, across scenarios (H-mode, L-mode), injected species, and injection locations. The measurements presented have formed the basis for systematic exhaust control on the considered devices. We now present this database for the essential validation of dynamic exhaust models for reactor design and control.

control↗

An Investigation of Heuristic Control Strategies for Multi-Electrolyzer Wind-Hydrogen Systems Considering Degradation

With the growing demand for renewable-energy-powered hydrogen generation and the corresponding increase in plant capacity, individually controlling many electrolyzer stacks will be critical for increasing the plant's lifetime and efficiency. This paper introduces a rule-based controller framework targeting electrolyzer degradation to explore the opportunity space in multi-stack hydrogen plant control. A novel control-oriented degradation model is also presented in this work, which quantifies the impact of steady power input, fluctuating power input, and the number of startup/shutdown cycles on electrolyzer degradation. Using a wind power input signal, 13 controller configurations are tested on 5 MW hydrogen plants consisting of 2, 5, 10, and 25 stacks. These configurations are evaluated on their capability to mitigate electrolyzer degradation and efficiently produce hydrogen from the time-varying input power. The results show that multi-stack control for degradation can extend the plant's lifetime by more than a factor of 3 with minimal impact on instantaneous hydrogen production.

control↗

Sink or Swim: A Tutorial on the Control of Floating Wind Turbines: Preprint

Within the rapidly growing wind energy sector, floating offshore wind turbines are expected to be the fastest growing portion. This is largely driven by the immense offshore wind resources that are mostly over deep water, where fixed-bottom concepts become cost prohibitive. However, compared to fixed-bottom wind turbines, floating wind turbines are more dynamic and exhibit potential instabilities, which requires advanced control technologies to ensure a safe and efficient operation. Beyond their existing objectives of maximizing power production while minimizing structural loads, floating wind turbine controllers must also avoid large platform oscillations and accommodate wave disturbances. This paper provides an overview of the challenges and opportunities in the control of floating offshore wind energy systems.

control↗