Search NASA⌕ Search

SEARCH · Search NASA

Results for “Approximate dynamic programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Configuring Airspace Sectors with Approximate Dynamic Programming

In response to changing traffic and staffing conditions, supervisors dynamically configure airspace sectors by assigning them to control positions. A finite horizon airspace sector configuration problem models this supervisor decision. The problem is to select an airspace configuration at each time step while considering a workload cost, a reconfiguration cost, and a constraint on the number of control positions at each time step. Three algorithms for this problem are proposed and evaluated: a myopic heuristic, an exact dynamic programming algorithm, and a rollouts approximate dynamic programming algorithm. On problem instances from current operations with only dozens of possible configurations, an exact dynamic programming solution gives the optimal cost value. The rollouts algorithm achieves costs within 2% of optimal for these instances, on average. For larger problem instances that are representative of future operations and have thousands of possible configurations, excessive computation time prohibits the use of exact dynamic programming. On such problem instances, the rollouts algorithm reduces the cost achieved by the heuristic by more than 15% on average with an acceptable computation time.

Bloem, Michael↗

A General Framework for Bounding Approximate Dynamic Programming Schemes

For years, there has been interest in approximation methods for solving dynamic programming problems, because of the inherent complexity in computing optimal solutions characterized by Bellman’s principle of optimality. A wide range of approximate dynamic programming (ADP) methods now exists. It is of great interest to guarantee that the performance of an ADP scheme be at least some known fraction, say ß , of optimal. This letter introduces a general approach to bounding the performance of ADP methods, in this sense, in the stochastic setting. The approach is based on new results for bounding greedy solutions in string optimization problems, where one has to choose a string (ordered set) of actions to maximize an objective function. This bounding technique is inspired by submodularity theory, but submodularity is not required for establishing bounds. Instead, the bounding is based on quantifying certain notions of curvature of string functions; the smaller the curvatures the better the bound. The key insight is that any ADP scheme is a greedy scheme for some surrogate string objective function that coincides in its optimal solution and value with those of the original optimal control problem. The ADP scheme then yields to the bounding technique mentioned above, and the curvatures of the surrogate objective determine the value ß of the bound. The surrogate objective and its curvatures depend on the specific ADP.

discrete event systems↗

Approximate Dynamic Programming With Enhanced Off-Policy Learning for Coordinating Distributed Energy Resources

Herein this paper proposes an innovative approximate dynamic programming (ADP) method for distributed energy resource coordination with the loss of life of battery energy storage system (BESS) explicitly modeled. The dispatch policy is designed to account for both calendrical and cyclical aging effects on BESS, explicitly modeling the impacts of ambient temperature on BESS lifespan. The proposed ADP employs an adaptive critic method and enhanced off-policy deterministic policy gradient (DPG) strategy, addressing the limitations of the on-policy gradient-based ADP approaches, including inadequate exploration, low data usage, and computational complexity. In particular, a customized policy is proposed to guide the algorithm to explore some promising decisions and thereby improve exploration capability and learning efficiency compared to conventional DPG-based learning approaches, which may struggle to find a global optimum due to random noisy action-based exploration or require expert demonstration with extra effort. The proposed method is illustrated using the IEEE 123-node system and compared with the existing ADP methods to prove solution accuracy and demonstrate the effects of incorporating degradation models into control design. Case studies showed that the proposed ADP effectively coordinates DERs with a 10 times smaller optimization gap compared to existing methods, and the incorporation of the BESS life loss model ensures the expected lifespan and results in significant cost savings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Experimental Validation of Approximate Dynamic Programming Based Optimization and Convergence on Microgrid Applications

Stochastic optimization can better address uncertainties in power system problems. However, when state space and action space become large, many existing approaches become computationally expensive and even infeasible. Approximate dynamic programming (ADP) attracts researchers’ attention as a powerful tool for solving power system optimization problems with reduced computational cost. In this paper, in light of the existing literature, we investigate how the ADP approach with post-decision value function approximation converges to the nearly optimal solution with improved computational speed and experimentally validate the performance of the approach for a microgrid energy optimization problem. The approximation error versus the number of iteration is studied for convergence analysis of the post-decision ADP. A flowchart is provided to illustrate the proposed ADP algorithm for a microgrid energy optimization problem. The performance of ADP and dynamic programming (DP) is compared in terms of optimization error and computational time. It has found that the post-decision ADP approach can achieve competitive optimality with improved computational speed compared to the traditional DP.

Das, Avijit↗

Real-Time Ecodriving Control in Electrified Connected and Autonomous Vehicles Using Approximate Dynamic Programing

Connected and automated vehicles (CAVs), particularly those with a hybrid electric powertrain, have the potential to significantly improve vehicle energy savings in real-world driving conditions. In particular, the ecodriving problem seeks to design optimal speed and power usage profiles based on available information from connectivity and advanced mapping features to minimize the fuel consumption over an itinerary. This paper presents a hierarchical multilayer model predictive control (MPC) approach for improving the fuel economy of a 48 V mild-hybrid powertrain in a connected vehicle environment. Approximate dynamic programing (DP) is used to solve the receding horizon optimal control problem, whose terminal cost is approximated with the base policy obtained from the long-term optimization. The controller was tested virtually (with deterministic and Monte Carlo simulation) across multiple real-world routes, demonstrating energy savings of more than 20%. The controller was then deployed on a test vehicle equipped with a rapid prototyping embedded controller. In-vehicle testing confirm the energy savings obtained in simulation and demonstrate the real-time ability of the controller.

Automation & Control Systems↗

Data-based and secure switched cyber–physical systems

In this work, we develop a completely model-free moving target defense framework for the detection and mitigation of sensor and/or actuator attacks in cyber–physical systems with dynamics that evolve in discrete-time. We incorporate an intrusion detection mechanism based on an approximate dynamic programming technique that learns the policies for optimal regulation and optimal tracking while simultaneously defending against actuator and sensor attacks in a model-free fashion. Switching rules are leveraged to force proactive and reactive defense mechanisms as well as, guarantee the stability of the equilibrium point. Finally, as a case study, we apply the proposed moving target defense framework to a DC–DC converter that is used in electric vehicles.

42 ENGINEERING↗

Integrated Optimization of Powertrain Energy Management and Vehicle Motion Control for Autonomous Hybrid Electric Vehicles

Hybrid Electric Vehicles (HEVs) and autonomous vehicles have been widely studied recently for on-road transportation. In the study of autonomous HEVs, the control of the vehicle's external dynamics and powertrain dynamics are often treated separately. Optimizing these two problems together can significantly improve fuel economy. In this paper, an autonomous HEV following a leader is considered. First, the augmented model to integrate the abovementioned dynamics is presented. Second, the optimization problem is defined to find the optimum fuel consumption of the follower in pursuit of a leader in a drive cycle. A customized control strategy based on Approximate Dynamic Programming (ADP) is then explored in which the optimal cost-to-go at each time step is approximated using neural networks. Also, the accuracy of the optimization solution is enhanced by applying the concept of the reachable sets. At last, three case studies show that the examined integrated control strategy outperforms the one with the separated optimization method by an additional 7.4%, 4.6%, and 11.8% improvement in fuel consumption, respectively.

33 ADVANCED PROPULSION SYSTEMS↗

Optimal, centralized dynamic curbside parking space zoning

In this paper we formulate a dynamic mixed integer program for optimally zoning curbside parking spaces subject to transportation policy-inspired constraints and regularization terms. First, we illustrate how given some objective of curb zoning valuation as a function of zone type (paid parking, bus stop, etc.), dynamically rezoning involves unrolling this optimization program over a fixed time horizon. Second, we implement two different solution methods given an example curb zoning valuation. In the first method, we solve long horizon dynamic zoning problems via approximate dynamic programming. In the second method, we employ Dantzig-Wolfe decomposition to break-up the mixed-integer program into a master problem and several sub-problems that can be solved in parallel. This speeds up the computational solve-time of the MIP considerably. We present simulation results and comparisons of the different employed techniques on vehicle arrival-rate data obtained for a neighborhood in downtown Seattle, Washington, USA.

Nazir, Mohammad Nawaf↗

Optimal Coordination of Distributed Energy Resources Using Deep Deterministic Policy Gradient

Recent studies showed that reinforcement learning (RL) is a promising approach for coordination and control of distributed energy resources (DER) under uncertainties. Many existing RL approaches, including Q-learning and approximate dynamic programming, are based on lookup table methods, which become inefficient when the problem size is large and infeasible when continuous states and actions are involved. In addition, when modeling battery energy storage system (BESS), the loss of life is not reasonably considered into the decision-making process. This paper proposes an innovative deep RL method for DER coordination considering BESS degradation. The proposed deep RL is designed based on an adaptive actor-critic architecture and employs an off-policy deterministic policy gradient method for determining the dispatch operation that minimizes the operation cost and BESS life loss. Case studies were performed to validate the proposed method and demonstrate the effects of incorporating degradation models into control design.

Das, Avijit↗

A multi-item maintenance center inventory model for low-demand reparable items

In many military and commercial contexts, complex equipment undergoes scheduled maintenance overhauls at regular intervals during which all failed components are replaced. Failure to have replacements on hand for all failed parts requires emergency measures at premium cost. When reparable parts are highly reliable and expensive, both holding and shortage costs are high. This model determines the reparable parts inventory for a maintenance center under three alternative criteria: (1) maximizing job-completion rate subject to constraint on total holding costs, (2) minimizing total holding costs plus expected job noncompletion costs, and (3) minimizing total holding costs subject to a required minimum job-completion rate. Exact solutions may be obtained using dynamic programming. Approximate solutions, found easily by marginal analysis, have readily computed bounds on possible error. The solution methods for the three formulations are illustrated in a simple example.

Schaefer, M. K.↗

Adaptive critic design-based reinforcement learning approach in controlling virtual inertia-based grid-connected inverters

In this report, an adaptive critic design (ACD) approach is proposed to control the phase and voltage of a grid-connected virtual synchronous generator (VSG). The penetration of fast responding inertia-less power converters significantly affect the stability of the power system, especially weak systems such as micro grids. The concept of virtual inertia addresses this concern by virtually emulating the behavior of a synchronous generator. However, the conventional VSG is designed based on two conditions: (i) fixed operating point and (ii) inductive grid connections. The performance of VSGs in low-voltage semi-resistive microgrids is far from optimal. To overcome the aforementioned concerns, a heuristic dynamic programing (HDP) approach is proposed to optimally control grid-connected VSGs. The neural-network-based inherence of the HDP enables the proposed technique to adapt to any impedance angle. The HDP controller includes two subnetworks: (i) the action network that controls the system optimally and (ii) the critic network, which evaluates the effectiveness of the action network. The simulation and experimental results are provided to evaluate the effectiveness of the proposed technique. As shown, the HDP-based approach illustrates a better performance in comparison with the conventional PI-based VSG in various operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

How good are learning-based control v.s. model-based control for load shifting? Investigations on a single zone building energy system

Both model predictive control (MPC) and deep reinforcement learning control (DRL) have been presented as a way to approximate the true optimality of a dynamic programming problem, and these two have shown significant operational cost saving potentials for building energy systems. Furthermore, there is still a lack of in-depth quantitative studies on their approximation levels to the true optimality, especially in the building energy domain. To fill in the gap, this paper provides a numerical framework that enables the evaluation of the optimality levels of different controllers for building energy systems. This framework is then used to comprehensively compare the optimal control performance of both MPC and DRL controllers with given computation budgets for a single zone fan coil unit system. Note the optimality is estimated based on a user-specific selection of trade-off weights among energy costs, thermal comfort and control slew rates. Compared with the best optimality we can find through expensive optimization simulations, the best DRL agent can maximally approximate the optimality by 96.54%, which outperforms the best MPC whose optimality level is 90.11%. However, due to the stochasticity, the DRL agent is only expected to approximate the optimality by 90.42%, which is almost equivalent to the best MPC. Except for Proximal Policy Optimization (PPO), all DRL agents can have a better approximation to the optimality than the best MPC, and are expected to have better approximation than the MPC with a prediction horizon of 32 steps (15 min per step). In terms of reducing energy cost and thermal discomfort, MPC can outperform the rule-based control (RBC) by 18.47%–25.44%. DRL can be expected to outperform RBC by 18.95%–25.65% ,and the best DRL control policy can outperform RBC by 20.29%–29.72%. Although the comparison of the optimality level is performed in a perfect setting, e.g., MPC assumes perfect models, and DRL assumes a perfect offline training process and online deployment process, this can shed insight on their capabilities of approximating to the original dynamic programming problem.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

42 ENGINEERING↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

Online anomaly detection↗

Control of Fractional Diffusion Problems via Dynamic Programming Equations

In this study, we explore the approximation of feedback control of integro-differential equations containing a fractional Laplacian term. To obtain feedback control for the state variable of this nonlocal equation, we use the Hamilton–Jacobi–Bellman equation. It is well known that this approach suffers from the curse of dimensionality, and to mitigate this problem we couple semi-Lagrangian schemes for the discretization of the dynamic programming principle with the use of Shepard approximation. This coupling enables approximation of high-dimensional problems. Numerical convergence toward the solution of the continuous problem is provided together with linear and nonlinear examples. The robustness of the method with respect to disturbances of the system is illustrated by comparisons with an open-loop control approach.

97 MATHEMATICS AND COMPUTING↗