Search NASA⌕ Search

SEARCH · Search NASA

Results for “Differentiable predictive control”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees

We present differentiable predictive control (DPC), a method for offline learning of constrained neural control policies for nonlinear dynamical systems with performance guarantees. We show that the sensitivities of the parametric optimal control problem can be used to obtain direct policy gradients. Specifically, we employ automatic differentiation (AD) to efficiently compute the sensitivities of the model predictive control (MPC) objective function and constraints penalties. To guarantee safety upon deployment, we derive probabilistic guarantees on closed-loop stability and constraint satisfaction based on indicator functions and Hoeffding’s inequality. We empirically demonstrate that the proposed method can learn neural control policies for various parametric optimal control tasks. In particular, we show that the proposed DPC method can stabilize systems with unstable dynamics, track time-varying references, and satisfy nonlinear state and input constraints. Our DPC method has practical time savings compared to alternative approaches for fast and memory-efficient controller design. Specifically, DPC does not depend on a supervisory controller as opposed to approximate MPC based on imitation learning. We demonstrate that, without losing performance, DPC is scalable with greatly reduced demands on memory and computation compared to implicit and explicit MPC while being more sample efficient than model-free reinforcement learning (RL) algorithms.

97 MATHEMATICS AND COMPUTING↗

Neural Lyapunov Differentiable Predictive Control

We present a learning-based predictive control methodology using the differentiable programming framework with sufficient Lyapunov-based guarantees. The neural Lyapunov differentiable predictive control (NLDPC) learns the policy by constructing a computational graph encompassing the system dynamics, state, and input constraints, and the necessary Lyapunov certification constraints, and thereafter using the automatic differentiation to update the neural policy parameters. In conjunction, our approach jointly learns a Lyapunov function that certifies the regions of state-space with stable dynamics. We also provide an initial condition sampling-based statistical guarantee for the training of NLDPC. Our offline training approach provides a computationally efficient and scalable alternative to classical model predictive control solutions, which now directly comes with Lyapunov based guarantees as shown in this paper. We substantiate the advantages of the proposed approach with simulations to stabilize the double integrator model and on an example of controlling an aircraft model.

Mukherjee, Sayak↗

Differentiable Predictive Control with Safety Guarantees: A Control Barrier Function Approach

In this paper, we develop a novel form of differentiable predictive control (DPC) with safety and robustness guarantees. DPC is a form of approximate model predictive control (MPC), wherein the control policy is a neural network that learns a receding horizon, optimal control law. The proposed approach exploits a new form of sampled-data barrier function to enforce safety, while only interrupting the neural network-based controller near the boundary of the safe set. The effectiveness of the proposed approach is demonstrated in simulation.

Shaw Cortez, Wenceslao E.↗

Learning Stochastic Parametric Differentiable Predictive Control Policies

We present a scalable unsupervised learning-based method for obtaining explicit control policies for model predictive control problems for stochastic linear systems with additive uncertainties subject to nonlinear chance constraints. We call the proposed method stochastic parametric differentiable predictive control (SP-DPC), which extends the recently proposed deterministic DPC policy optimization algorithm. We formulate the SP-DPC as a deterministic approximation to the stochastic parametric constrained optimal control problem via independent sampling of the problem's parameters and uncertainties. This formulation allows us to directly compute the policy gradients via automatic differentiation of the problem's value function, evaluated over sampled parameters and uncertainties. In particular, the computed expectation of the problem's value function is backpropagated through the finite-time closed-loop system rollouts parametrized by a known nominal system dynamics model and neural control policy. We also provide theoretical probabilistic guarantees on closed-loop stability and chance constraints satisfaction for systems controlled by learned neural policies. We demonstrate the computational efficiency and scalability of the proposed policy optimization algorithm in three numerical examples, including systems with a large number of states or subject to nonlinear constraints.

Drgona, Jan↗

Koopman-based Differentiable Predictive Control for the Dynamics-Aware Economic Dispatch Problem

The dynamics-aware economic dispatch (DED) problem embeds low-level generator dynamics and operational constraints to enable near real-time scheduling of generation units in a power network. DED produces a more dynamic supervisory control policy than traditional economic dispatch (T-ED) that reduces overall generation costs. However, the incorporation of differential equations that govern the system dynamics makes DED an optimization problem that is computationally prohibitive to solve. In this work, we present a new data-driven approach based on differentiable programming to efficiently obtain offline parametric solutions to the underlying DED problem. In particular, we employ the recently proposed differentiable predictive control (DPC) for offline learning of explicit neural control policies based on identified Koopman operator (KO) model of the system dynamics. We demonstrate the high solution quality and five orders of magnitude computational-time savings of the DPC method over the original optimization-based DED approach on a 9-bus test power grid network.

King, Ethan↗

LBC (Learning Building Control)

LBC encompasses the source code and data to reproduce and extend results for a manuscript that compares demand responsive control schemes for multi-zone buildings. It includes examples of model predictive control (MPC), value function approximation via CVXPYLAYERS, and differentiable predictive control (DPC). The goal of the study is to evaluate state-of-the-art controllers and establish the efficacy (if any) of learning-based approaches that leverage deep neural networks in one way or another.

Zhang, Xiangyu↗

Dynamic Decarbonization through Autonomous Physics-Centric Deep Learning and Optimization of Building Operations (Abstract only)

This project directly addresses the primary goal of Area of Interest 2 in the CRADA call: to advance optimization-based integrated energy management systems in commercial and residential buildings. Pacific Northwest National Laboratory (PNNL) and its industry partner PassiveLogic aim to accomplish this by reaching three key objectives. First, to ensure a broad impact in the building controls industry, PNNL will extend its open-source library for predictive control synthesis by augmenting its capabilities with data-driven self-learning of building models and auto-calibration of predictive controllers. The effort will focus on building use cases selected in collaboration with PassiveLogic. The team will specifically address the development of methods for data-driven adaptation of building models, investigation of model architectures that best address specific building types, and automated synthesis of differentiable predictive controllers that optimize diverse objectives. Second, PNNL will collaborate with PassiveLogic to integrate the aforementioned methods with PasiveLogic’s advanced controls platform. The collaborative integration effort will inform the developments under the first objective by providing specific data on the attainable performance of model learning on resource-constrained edge computing platforms. This software integration effort will increase the technical maturity of the developed libraries by exploring the use of software integration tools and methods. Third, PNNL and PassiveLogic will work to improve the technology readiness of the developed predictive controllers by testing their performance in relevant test environments, such as high-fidelity simulation, hardware in the loop, and actual test buildings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Path Planning: Differential Dynamic Programming and Model Predictive Path Integral Control on VTOL Aircraft

This paper explores two optimal control approaches, widely used in robotics, to establish their viability as real-time trajectory planners for vehicle configurations envisioned for the emerging aviation sector of Urban Air Mobility (UAM). Differential Dynamic Programming (DDP) enables planning over highly nonlinear dynamics using second-order approximations along a nominal trajectory, and displays quadratic convergence to a local solution. Model Predictive Path Integral (MPPI) is a stochastic sampling-based algorithm that can optimize for general cost criteria, including potentially highly nonlinear formulations, and supports parallel computation through the use of modern GPU hardware. In this work, DDP and MPPI were implemented using model predictive control (MPC), and the results indicate they are able to successfully transition the aircraft over different flight envelopes and generate trajectories unique to UAM vehicles.

Differential Dynamic Programming↗

Preliminary Development of Heat Transfer Model-Based Control Algorithms of Liquid Sodium Purification System: Advanced Sensors and Instrumentation Advanced Controls

Monitoring the operation of sodium purification system is essential for efficient operation of sodium fast reactors. In this work, a heat transfer model has been developed for monitoring the plugging meter and cold trap systems at the Mechanisms Engineering Test Loop (METL) liquid sodium facility at Argonne National Laboratory. The model of the purification system was developed by treating the respective aspects of the cold trap purification loop and plugging meter diagnostic loop as two separate control volumes using information from the METL piping and instrumentation diagram (P&ID). A model predictive controller was designed using first order differential equations with the specified boundary conditions. The system behavior was studied with a tuned optimized procedure using the internal cold trap temperature and plugging meter outlet temperature as control variables, and the air blower temperature as an independent variable respectively. Results of computer simulations obtained in this study compared favorably with experimental data showing very good reference tracking response with negligible overshoot as both plugging meter and cold trap physical models approach the setpoint.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Adaptive Control of a Transport Aircraft Using Differential Thrust

The paper presents an adaptive control technique for a damaged large transport aircraft subject to unknown atmospheric disturbances such as wind gust or turbulence. It is assumed that the damage results in vertical tail loss with no rudder authority, which is replaced with a differential thrust input. The proposed technique uses the adaptive prediction based control design in conjunction with the time scale separation principle, based on the singular perturbation theory. The application of later is necessitated by the fact that the engine response to a throttle command is substantially slow that the angular rate dynamics of the aircraft. It is shown that this control technique guarantees the stability of the closed-loop system and the tracking of a given reference model. The simulation example shows the benefits of the approach.

Stepanyan, Vahram↗

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics↗

Control theory for random systems

A survey is presented of the current knowledge available for designing and predicting the effectiveness of controllers for dynamic systems which can be modeled by ordinary differential equations. A short discussion of feedback control is followed by a description of deterministic controller design and the concept of system state. The need for more realistic disturbance models led to the use of stochastic process concepts, in particular the Gauss-Markov process. A compensator controlled system, with random forcing functions, random errors in the measurements, and random initial conditions, is treated as constituting a Gauss-Markov random process; hence the mean-square behavior of the controlled system is readily predicted. As an example, a compensator is designed for a helicopter to maintain it in hover in a gusty wind over a point on the ground.

Bryson, A. E., Jr.↗