Search NASA⌕ Search

SEARCH · Search NASA

Results for “Markov decision processes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Decision-making based on Markov decision process in integrated artificial reasoning framework—Part I: Theory

This paper presents a decision-making framework based on an integrated artificial reasoning framework and Markov decision process (MDP). The integrated artificial reasoning framework provides a physics-based approach that converts system information into state transition models, and the analysis result will be represented by the transition probabilities that can be used with an MDP to find a traceable and explainable optimal pathway. A dynamic Bayesian network (DBN) is well suited for representing the structure of an MDP. The causality information among process variables (or among subsystems) is mathematically represented in a DBN by the conditional probabilities of the node’s states provided different probabilities of the parent node’s states. To define node states in a physically understandable manner, we used multilevel flow modeling (MFM). An MFM follows the fundamental energy and mass conservation laws and supports the selection of process variables that represent the system of interest so that causal relations among process variables are properly captured. An MFM-based DBN supports developing state transition models in an MDP to capture the effect of process variables of system having physical relations. The operators of the target system can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from component degradation or random failures. We analyzed a simplified exemplary system to illustrate an optimal operational policy using the suggested approach.

Markov decision process↗

Quantum logic gate synthesis as a Markov decision process

Reinforcement learning has witnessed recent applications to a variety of tasks in quantum programming. The underlying assumption is that those tasks could be modeled as Markov decision processes (MDPs). Here, we investigate the feasibility of this assumption by exploring its consequences for single-qubit quantum state preparation and gate compilation. By forming discrete MDPs, we solve for the optimal policy exactly through policy iteration. We find optimal paths that correspond to the shortest possible sequence of gates to prepare a state or compile a gate, up to some target accuracy. Our method works in both the absence and presence of noise and compares favorably to other quantum compilation methods, such as the Ross–Selinger algorithm. This work provides theoretical insight into why reinforcement learning may be successfully used to find optimally short gate sequences in quantum programming.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Markov Decision Processes for Intelligent, Risk-Informed Asset-Management Decision-Making

Advanced nuclear reactors are a promising option for aiding the world in achieving its net-zero carbon emission goals, however, there are significant challenges to attaining and maintaining economic competitiveness with other sources of electricity. To improve the economic competitiveness of advanced reactor designs, a project was initiated to explore the use of Markov Decision Processes (MDPs) to guide asset-management decision-making during advanced reactor operation. MDPs are a powerful tool for optimizing decision-making in complex environments and their application to advanced reactors can aid in planning maintenance and repair activities to minimize downtime and maximize generation. The described approach expands on previous work regarding the use of MDPs for operational decision-making through the direct incorporation of real-time plant information. The integral MDP analysis includes information from online component diagnostic tools and the plant’s real-time generation risk assessment (GRA) and probabilistic risk assessment (PRA), which evaluate plant risk from both an economic and safety perspective. The result is an asset-management optimization framework that is based on real-time data regarding plant component status and the current best-estimate of plant risk. The paper presents an overview of the theoretical framework to incorporate the different information pathways into an integral MDP analysis, along with example analyses.

Grabaskas, David↗

Operation Optimization using Reinforcement Learning with Integrated Artificial Reasoning Framework

In large and complex systems, operational decision-making requires a systematic analysis with a vast amount of data from both process parameters and component status monitoring. In this paper, we present an integrated artificial reasoning approach for system state transition models that can help operational decision-making with explainable and traceable reasoning. The integrated artificial reasoning framework is a physics-based approach of defining the system structure in a Bayesian network, so we leveraged it in a Markov decision process (MDP) for finding optimal operational solutions. In our proposed framework, the MDP is implemented on a dynamic Bayesian network (DBN), which represents causalities in a system. The multilevel flow modeling was utilized in order to extract these causalities in a more efficient and objective manner. Since multilevel flow modeling is based on the fundamental energy and mass conservation laws, the target system is decomposed into several mass, energy, and information structures, which serve as the basis for a DBN. The MDP consists of the processes of finding a solution for the Bellman equation, which can be derived from the conditional probability equations of the constructed DBN. System operators can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from the component degradation process or random failures. We analyzed a simplified example system to illustrate finding an optimal operational policy with this approach.

99 GENERAL AND MISCELLANEOUS↗

Deep Reinforcement Learning for Distribution System Operations: A Tutorial and Survey

Here, the rapid evolution of modern electric power distribution systems into complex networks of interconnected active devices, distributed generation (DG), and storage poses increasing difficulties for system operators. The large-scale integration of distributed energy resources (DERs) and the rapid exchange of measurement data via communication networks present major opportunities for advancing grid operations but also introduce greater uncertainty, higher data dimensionality, more complex network and device models, and challenging control and optimization problems. Deep reinforcement learning (DRL) algorithms are promising in addressing these challenges. However, they have not been effectively adapted for power systems applications, requiring extensive customization for implementation and evaluation. This has resulted in reproducibility challenges and a steep learning curve for researchers new to applying DRL algorithms to the power systems domain. To bridge these gaps, this tutorial aims to serve as a valuable resource for researchers interested in exploring learning-based algorithms to operate active power distribution networks. Specifically, this work presents a generalized process for translating sequential decision-making problems in power distribution systems into Markov decision process (MDP) formulations, illustrated through concrete grid service examples. Additionally, we introduce a simple environment design strategy to develop and evaluate example DRL algorithms for distribution system applications, complete with an included code repository to guide users through environment construction.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reinforcement Learning for Intentional Islanding in Resilient Power Transmission Systems

Intentional islanding is the process of identifying and deliberately decomposing the transmission network to form self-sustained islands from an endangered network during disruptions to improve resilience and security. Most existing intentional islanding models are offline resilience decision tools and hence do not provide outage responses in a timely manner. In this paper, a reinforcement learning (RL) based model for intentional islanding is developed, which offers real-time switching control, online deployability, and adaptability to varying system conditions. The intentional islanding process is formulated as a Markov decision process, where the optimal transmission switching policy is learned using the RL approach. The control policy is learned over an environment that encompasses a Power System Simulator for Engineering (PSS/E) model of the transmission network, facilitated by an interface to the standard openAI Gym framework. The proposed RL-based methodology aims to form stable and self-sustainable islands by ensuring voltage stability while reducing the power mismatch in the formed islands. A proximal policy optimization algorithm is designed, which is suitable for controlling the on/off status of the switches with multi-layer perceptron as value and actor networks. The effectiveness of the proposed framework in the self-recovery of the grid by island formation is applied on the modified IEEE 39-bus test network and validated by dynamic simulations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Bayesian sequential optimal experimental design for nonlinear models using policy gradient reinforcement learning

We present a mathematical framework and computational methods for optimally designing a finite sequence of experiments. This sequential optimal experimental design (sOED) problem is formulated as a finite-horizon partially observable Markov decision process (POMDP) under a Bayesian setting and with information-theoretic utilities. The formulation is general and may accommodate continuous random variables, non-Gaussian posteriors, and nonlinear forward models. The sOED design policy incorporates elements of feedback and lookahead simultaneously, and we show it to generalize the commonly-used batch and greedy design strategies. We solve for the sOED policy using the policy gradient (PG) method from reinforcement learning, and provide a derivation for the PG expression in the sOED context. Adopting an actor-critic approach, the policy and value functions are parameterized using deep neural networks and improved via PG estimates produced from simulated episodes of designs and observations. The new PG-sOED algorithm is first validated on a linear-Gaussian benchmark, and then compared against other design baselines on a sensor movement problem for contaminant source inversion in a convection-diffusion field. As a result, we provide explanation for the policy behaviors using knowledge of the underlying physical process.

97 MATHEMATICS AND COMPUTING↗

Network Reconfiguration for Enhanced Operational Resilience Using Reinforcement Learning

This paper proposes a reinforcement learning-based approach for distribution network reconfiguration(DNR) to enhance the resilience of the electric power supply. Resilience enhancements usually require solving large-scale stochastic optimization problems that are computationally expensive and sometimes infeasible. The exceptional performance of reinforcement learning techniques has encouraged their adoption in various power system control studies, specifically resilience-based real-time applications. In this paper, a single agent framework is developed using an Actor-Critic algorithm (ACA) to determine statuses of tie-switches in a distribution feeder impacted by an extreme weather event. The proposed approach provides a fast-acting control algorithm that reconfigures the feeder topology to reduce or even avoid load shedding. The problem is formulated as a discrete Markov decision process in such a way that a system state captures the system topology and its operational characteristics. An action is made to open or close a specific set of tie-switches after which a reward is calculated to evaluate the practicality and advantage of that action. The iterative Markov process is used to train the proposed ACA under diverse failure scenarios and is demonstrated on the 33-node distribution feeder system. Results show the capability of the proposed ACA to determine proper switching action of tie-switches with accuracy exceeding 93%.

actor critic↗

Deep Reinforcement Learning for Distribution System Restoration Using Distributed Energy Resources and Tie-Switches

Distributed energy resources (DERs), such as solar PVs and energy storage, can be used to restore distribution system critical loads after the extreme weather events to increase grid resilience. However, coordinating multiple DERs together with tie-switches for multi-step restoration process under renewable uncertainty is challenging. This paper proposes a deep reinforcement learning to control discrete actions of switching on/off tie switches and DERs for critical load restoration. The restoration problem is first cast into the Markov decision process suitable for DRL. Then, the original soft actor critic (SAC) method for continuous actions has been extended to handle discrete and continuous actions. Numerical comparison results with other stochastic optimization-based approaches on the modified IEEE 33-bus system show that the proposed method can achieve fast critical load restoration in the presence of substation power outage while maintaining system voltage limit throughout the restoration process.

active distribution systems↗

Reinforcement Learning-Based Oscillation Dampening: Scaling Up Single-Agent Reinforcement Learning Algorithms to a 100-Autonomous-Vehicle Highway Field Operational Test

In this article, we explore the technical details of the reinforcement learning (RL) algorithms that were deployed in the largest field test of automated vehicles designed to smooth traffic flow in history as of 2023, uncovering the challenges and breakthroughs that come with developing RL controllers for automated vehicles. We delve into the fundamental concepts behind RL algorithms and their application in the context of self-driving cars, discussing the developmental process from simulation to deployment in detail, from designing simulators to reward function shaping. We present the results in both simulation and deployment, discussing the flow-smoothing benefits of the RL controller. From understanding the basics of Markov decision processes to exploring advanced techniques such as deep RL, our article offers a comprehensive overview and deep dive of the theoretical foundations and practical implementations driving this rapidly evolving field. We also showcase real-world case studies and alternative research projects that highlight the impact of RL controllers in revolutionizing autonomous driving. From tackling complex urban environments to dealing with unpredictable traffic scenarios, these intelligent controllers are pushing the boundaries of what automated vehicles can achieve. Furthermore, we examine the safety considerations and hardware-focused technical details surrounding deployment of RL controllers into automated vehicles. As these algorithms learn and evolve through interactions with the environment, ensuring their behavior aligns with safety standards becomes crucial. Here, we explore the methodologies and frameworks being developed to address these challenges, emphasizing the importance of building reliable control systems for automated vehicles.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

"Reinforcement Learning Based Oscillation Dampening: Scaling up Single-Agent RL algorithms to a 100 AV highway field operational test"

In this article, we explore the technical details of the reinforcement learning (RL) algorithms that were deployed in the largest field test of automated vehicles designed to smooth traffic flow in history as of 2023, uncovering the challenges and breakthroughs that come with developing RL controllers for automated vehicles. We delve into the fundamental concepts behind RL algorithms and their application in the context of self-driving cars, discussing the developmental process from simulation to deployment in detail, from designing simulators to reward function shaping. We present the results in both simulation and deployment, discussing the flow-smoothing benefits of the RL controller. From understanding the basics of Markov decision processes to exploring advanced techniques such as deep RL, our article offers a comprehensive overview and deep dive of the theoretical foundations and practical implementations driving this rapidly evolving field. We also showcase real-world case studies and alternative research projects that highlight the impact of RL controllers in revolutionizing autonomous driving. From tackling complex urban environments to dealing with unpredictable traffic scenarios, these intelligent controllers are pushing the boundaries of what automated vehicles can achieve. Furthermore, we examine the safety considerations and hardware-focused technical details surrounding deployment of RL controllers into automated vehicles. As these algorithms learn and evolve through interactions with the environment, ensuring their behavior aligns with safety standards becomes crucial. We explore the methodologies and frameworks being developed to address these challenges, emphasizing the importance of building reliable control systems for automated vehicles.

Jang, Kathy↗

Beyond PID Controllers: PPO with Neuralized PID Policy for Proton Beam Intensity Control in Mu2e

We introduce a novel Proximal Policy Optimization (PPO) algorithm aimed at addressing the challenge of maintaining a uniform proton beam intensity delivery in the Muon to Electron Conversion Experiment (Mu2e) at Fermi National Accelerator Laboratory (Fermilab). Our primary objective is to regulate the spill process to ensure a consistent intensity profile, with the ultimate goal of creating an automated controller capable of providing real-time feedback and calibration of the Spill Regulation System (SRS) parameters on a millisecond timescale. We treat the Mu2e accelerator system as a Markov Decision Process suitable for Reinforcement Learning (RL), utilizing PPO to reduce bias and enhance training stability. A key innovation in our approach is the integration of a neuralized Proportional-Integral-Derivative (PID) controller into the policy function, resulting in a significant improvement in the Spill Duty Factor (SDF) by 13.6%, surpassing the performance of the current PID controller baseline by an additional 1.6%. This paper presents the preliminary offline results based on a differentiable simulator of the Mu2e accelerator. It paves the groundwork for real-time implementations and applications, representing a crucial step towards automated proton beam intensity control for the Mu2e experiment.

43 PARTICLE ACCELERATORS↗

Reinforcement learning for real-time process control in high-temperature superconductor manufacturing

With high efficiency and low energy loss, high-temperature superconductors (HTS) have demonstrated their profound applications in various fields, such as medical imaging, transportation, accelerators, microwave devices, and power systems. The high-field applications of HTS tapes have raised the demand for producing cost-effective tapes with long lengths in superconductor manufacturing. However, achieving the uniform and enhanced performance of a long HTS tape is challenging due to the unstable growth conditions in the manufacturing process. Although it is confirmed that the process parameters during the advanced metal organic chemical vapor deposition (A-MOCVD) process influence the uniformity of the produced HTS tapes, the high-dimensional process parameter signals and their complicated interactions make it difficult to develop an effective control policy. In this paper, we propose a local measure for the uniformity of HTS tapes to provide instant feedback for our control policy. Then, we model the manufacturing of HTS tapes as a Markov decision process (MDP) with continuous state and action spaces to assess the instant reward in real time in our feedback control model. As our MDP involves continuous and high-dimensional state and action spaces, a neural fitted Q-iteration (NFQ) algorithm is adopted to solve the MDP with artificial neural network (ANN) function approximation. The collinearity of process parameters can restrict our capability of adjusting the process parameters, which is addressed by the principal component analysis (PCA) in our method. The control policy adjusts the PCA of process parameters using the NFQ algorithm. In conclusion, based on our case studies on real A-MOCVD dataset, the obtained control policy increases the average uniformity of tapes by 5.6% and performs especially well on sample HTS tapes with a low uniformity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Operation Optimization Using Reinforcement Learning with Integrated Artificial Reasoning Framework

In large and complex systems, operational decision-making requires a systematic analysis with a vast amount of data from both process parameters and component status monitoring. In this paper, we present an integrated artificial reasoning approach for system state transition models that can help operational decision-making with explainable and traceable reasoning. The integrated artificial reasoning framework is a physics-based approach of defining the system structure in a Bayesian network, so we leveraged it in a Markov decision process (MDP) for finding optimal operational solutions. In our proposed framework, the MDP is implemented on a dynamic Bayesian network (DBN), which represents causalities in a system. The multilevel flow modeling was utilized in order to extract these causalities in a more efficient and objective manner. Since multilevel flow modeling is based on the fundamental energy and mass conservation laws, the target system is decomposed into several mass, energy, and information structures, which serve as the basis for a DBN. The MDP consists of the processes of finding a solution for the Bellman equation, which can be derived from the conditional probability equations of the constructed DBN. System operators can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from the component degradation process or random failures. We analyzed a simplified example system to illustrate finding an optimal operational policy with this approach.

Kim, Junyung↗

GraMeR: Gra ph Me ta R einforcement learning for multi-objective influence maximization

Influence maximization (IM) is a combinatorial problem of identifying a subset of seed nodes in a network (graph), which when activated, provide a maximal spread of influence in the network for a given diffusion model and a budget for seed set size. IM has numerous applications such as viral marketing, epidemic control, sensor placement and other network-related tasks. However, its practical uses are limited due to the computational complexity of current algorithms. Recently, deep reinforcement learning has been leveraged to solve IM in order to ease the computational burden. However, there are serious limitations in current approaches, including narrow IM formulation that only consider influence via spread and ignore self-activation, low scalability to large graphs, and lack of generalizability across graph families leading to a large running time for every test network. In this work, we address these limitations through a unique approach that involves: (1) Formulating a generic IM problem as a Markov decision process that handles both intrinsic and influence activations; (2)incorporating generalizability via meta-learning across graph families. There are previous works that combine deep reinforcement learning with graph neural network, but this work solves a more realistic IM problem and incorporates generalizability across graphs via meta reinforcement learning. Extensive experiments are carried out in various standard networks to validate performance of the proposed Graph Meta Reinforcement learning (GraMeR) framework. Finally, the results indicate that GraMeR is multiple orders faster and generic than conventional approaches when applied on small to medium scale graphs.

97 MATHEMATICS AND COMPUTING↗

Autonomous Cyber Defense Against Dynamic Multi-strategy Infrastructural DDoS Attacks

Dynamic Infrastructural Distributed Denial of Service (I-DDoS) attacks constantly change attack vectors to congest core backhaul links and disrupt critical network availability while evading end-system defenses. To effectively counter these highly dynamic attacks, defense mechanisms need to exhibit adaptive decision strategies for real-time mitigation. This paper presents a novel Autonomous DDoS Defense framework that employs model-based reinforcement agents. The framework continuously learns attack strategies, predicts attack actions, and dynamically determines the optimal composition of defense tactics such as filtering, limiting, and rerouting for flow diversion. Our contributions include extending the underlying formulation of the Markov Decision Process (MDP) to address simultaneous DDoS attack and defense behavior, and accounting for environmental uncertainties. We also propose a fine-grained action mitigation approach robust to classification inaccuracies in Intrusion Detection Systems (IDS). Additionally, our reinforcement learning model demonstrates resilience against evasion and deceptive attacks. Evaluation experiments using real-world and simulated DDoS traces demonstrate that our autonomous defense framework ensures the delivery of approximately 96 - 98% of benign traffic despite the diverse range of attack strategies.

Dutta, Ashutosh↗

An Online Approach to Solve the Dynamic Vehicle Routing Problem with Stochastic Trip Requests for Paratransit Services

Many transit agencies operating paratransit and microtransit services have to respond to trip requests that arrive in real-time, which entails solving hard combinatorial and sequential decision-making problems under uncertainty. To avoid decisions that lead to significant inefficiency in the long term, vehicles should be allocated to requests by optimizing a non-myopic utility function or by batching requests together and optimizing a myopic utility function. While the former approach is typically offline, the latter can be performed online. We point out two major issues with such approaches when applied to paratransit services in practice. First, it is difficult to batch paratransit requests together as they are temporally sparse. Second, the environment in which transit agencies operate changes dynamically (e.g., traffic conditions can change over time), causing the estimates that are learned offline to become stale. To address these challenges, we propose a fully online approach to solve the dynamic vehicle routing problem (DVRP) with time windows and stochastic trip requests that is robust to changing environmental dynamics by construction. We focus on scenarios where requests are relatively sparse—our problem is motivated by applications to paratransit services. We formulate DVRP as a Markov decision process and use Monte Carlo tree search to evaluate actions for any given state. Accounting for stochastic requests while optimizing a non-myopic utility function is computationally challenging; indeed, the action space for such a problem is intractably large in practice. To tackle the large action space, we leverage the structure of the problem to design heuristics that can sample promising actions for the tree search. Our experiments using real-world data from our partner agency show that the proposed approach outperforms existing state-of-the-art approaches both in terms of performance and robustness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗