Search NASA⌕ Search

SEARCH · Search NASA

Results for “model-free control”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Model-Free Control of Grid-Interactive Efficient Buildings Under Communication Time Delays

Grid-interactive efficient buildings (GEBs) have recently been used to enhance the reliability and stability of the electric grid through demand response (DR) programs. However, most existing DR control strategies require accurate modeling of the various building thermostatically controlled loads (TCLs) and are computationally expensive. To address these challenges, a model-free control (MFC)-based strategy has recently been introduced for coordinating and controlling GEBs. MFC is a data-enabled control strategy that is computationally efficient and does not require the analytical models of the various building equipment. In this paper, we numerically investigate the impact of communication time delays on the performance of MFC in maintaining the TCLs' temperatures within the desired comfort levels while meeting the assigned power allocation constraint.

Telsang, Bhagyashri [University of Tennessee, Knox↗

Adaptive Model-Free Vehicle Path-Tracking via Fast-Converging Prescribed-Time Newton-Based Extremum-Seeking Control

Model-free control (MFC) offers a simple and effective approach to automated vehicle path-tracking without requiring an explicit plant model for control law design. However, gain tuning in MFC is typically carried out through trial-and-error, which can be time-consuming and may lead to suboptimal performance. To address this limitation, extremum-seeking-based adaptive MFC has shown promise by enabling real-time adaptation of control gains, without relying on a predefined vehicle model. Nonetheless, existing ESC approaches often suffer from slow convergence. This paper integrates MFC, employing longitudinal and lateral ultra-local models of a rear-wheel-drive vehicle, with a novel prescribed-time (PT) Newton-based extremum-seeking control (ESC) strategy that ensures rapid convergence of control gains within the prescribed time. Unlike conventional gradient-based ESC methods, the PT Newton-based ESC leverages artificial delays and time-periodic gains, not only to guarantee convergence within the specified time, but also to compensate for feedback delays. Simulation results demonstrate that the proposed approach significantly improves gain adaptation speed and tracking accuracy. This work advances adaptive model-free vehicle control by offering a high-performance, delay-resilient alternative to existing ESCMFC frameworks.

Waleed khan, Muhammad [The University of Texas at ↗

Cooperative On-Ramp Merging with Time-Varying Vehicle-to-Vehicle Communication Delay Compensation via a Model-Free Approach

Cooperative merging strategies enabled by vehicle-to-vehicle (V2V) communication have shown promise in addressing congestion, fuel inefficiency, and collision risks. However, their performance can be severely degraded by time-varying and uncertain communication delays-an issue often overlooked in existing research, which primarily focuses on merging sequence determination and trajectory planning. Furthermore, practical considerations such as heterogeneous vehicle dynamics, varying road conditions, and real-time implementation complexities are frequently neglected. This paper presents a model-free, online planning framework for cooperative on-ramp merging of connected and automated vehicles (CAVs), explicitly accounting for time-varying V2V communication delays. Without relying on detailed vehicle dynamics, the proposed method introduces a data-driven delay compensation scheme. A co-simulation platform integrating high-fidelity vehicle dynamics, traffic simulation (SUMO), and V2V communication within MATLAB/Simulink is developed to evaluate the proposed method. Simulation results demonstrate that unaddressed V2V communication delays significantly impair merging performance. In contrast, the proposed framework enhances intervehicle distance tracking and maintains low CO2 emissions and fuel consumption, under communication delay across different communication frequencies. In conclusion, its lightweight design also facilitates real-time implementation, making it well-suited for deployment in practical CAV systems.

Accounting↗

An adaptive model-free robotic force control strategy for hydrodynamic real-time hybrid simulation of floating offshore wind turbines

Real-time hybrid simulation (RTHS) - a cyber-physical testing approach - promises to enhance the simulation fidelity of the model-scale experiments used to prototype floating offshore wind turbines (FOWTs). In hydrodynamic RTHS (hydro-RTHS), actuators emulate aerodynamic forces on model-scale FOWT specimens subjected to physical waves in a hydrodynamic laboratory. Robotic arms are promising candidates for actuation in hydro-RTHS due to their compact multi-degree-of-freedom (DOF) capabilities. Unlike classical RTHS for seismic applications, which typically relies on displacement control, hydro-RTHS requires 6-DOF force control on newly designed floating prototypes in a model-scale setting, which presents significant challenges, including modeling uncertainties, directional asymmetry, configuration drift, bandwidth limitations, and time-varying delays. To mitigate these constraints without extensive pre-test calibration, this study proposes an adaptive model-free robotic force control strategy that combines task-space explicit force control with a secondary joint-space pose-keeping task. The Adaptive Feedforward Compensator (AFC) is integrated into the force control loop to compensate for time-varying delay. Experimental testing was conducted using a Franka Emika Panda robotic arm with a 1:50 scale FOWT specimen under operational wind and wave conditions. Results demonstrate stable and consistent 6-DOF force tracking. Effective delay compensation was observed, with low-frequency delay reductions ranging from 71.4% to 91.8% and improvements in low-frequency surge force tracking of 25.0% to 52.1%. This study enhances robotic actuation performance in hydro-RTHS and introduces a force control strategy that supports reliable robotic operation in uncertain floating environments. Future work will explore disturbance-observer mechanisms to further enhance wave rejection capabilities under extreme wind and wave conditions.

17 WIND ENERGY↗

A safe reinforcement learning algorithm for supervisory control of power plants

Traditional control theory-based methods require tailored engineering for each system and constant fine-tuning. In power plant control, one often needs to obtain a precise representation of the system dynamics and carefully design the control scheme accordingly. Model-free Reinforcement learning (RL) has emerged as a promising solution for control tasks due to its ability to learn from trial-and-error interactions with the environment. It eliminates the need for explicitly modeling the environment’s dynamics, which is potentially inaccurate. However, the direct imposition of state constraints in power plant control raises challenges for standard RL methods. To address this, we propose a chance-constrained RL algorithm based on Proximal Policy Optimization for supervisory control. Our method employs Lagrangian relaxation to convert the constrained optimization problem into an unconstrained objective, where trainable Lagrange multipliers enforce the state constraints. In conclusion, our approach achieves the smallest distance of violation and violation rate in a load-follow maneuver for an advanced Nuclear Power Plant design.

constrained optimization↗

Entanglement engineering of optomechanical systems by reinforcement learning

Entanglement is fundamental to quantum information science and technology, yet controlling and manipulating entanglement—so-called entanglement engineering—for arbitrary quantum systems remains a formidable challenge. There are two difficulties: the fragility of quantum entanglement and its experimental characterization. We develop a model-free deep reinforcement-learning (RL) approach to entanglement engineering, in which feedback control together with weak continuous measurement and partial state observation is exploited to generate and maintain desired entanglement. We employ quantum optomechanical systems with linear or nonlinear photon–phonon interactions to demonstrate the workings of our machine-learning-based entanglement engineering protocol. In particular, the RL agent sequentially interacts with one or multiple parallel quantum optomechanical environments, collects trajectories, and updates the policy to maximize the accumulated reward to create and stabilize quantum entanglement over an arbitrary amount of time. The machine-learning-based model-free control principle is applicable to the entanglement engineering of experimental quantum systems in general.

97 MATHEMATICS AND COMPUTING↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗

Efficient and assured reinforcement learning-based building HVAC control with heterogeneous expert-guided training

Abstract Building heating, ventilation, and air conditioning (HVAC) systems account for nearly half of building energy consumption and $$20\%$$ of total energy consumption in the US. Their operation is also crucial for ensuring the physical and mental health of building occupants. Compared with traditional model-based HVAC control methods, the recent model-free deep reinforcement learning (DRL) based methods have shown good performance while do not require the development of detailed and costly physical models. However, these model-free DRL approaches often suffer from long training time to reach a good performance, which is a major obstacle for their practical deployment. In this work, we present a systematic approach to accelerate online reinforcement learning for HVAC control by taking full advantage of the knowledge from domain experts in various forms . Specifically, the algorithm stages include learning expert functions from existing abstract physical models and from historical data via offline reinforcement learning, integrating the expert functions with rule-based guidelines, conducting training guided by the integrated expert function and performing policy initialization from distilled expert function. Moreover, to ensure that the learned DRL-based HVAC controller can effectively keep room temperature within the comfortable range for occupants, we design a runtime shielding framework to reduce the temperature violation rate and incorporate the learned controller into it. Experimental results demonstrate up to 8.8 X speedup in DRL training from our approach over previous methods, with low temperature violation rate.

Xu, Shichao↗

Time-Varying Output Delay Compensation-A Model-Free Approach and its Application on Cooperative On-Ramp Merging

This paper presents a model-free approach to compensate for time-varying output delay in networked control systems. The proposed architecture combines a model-free observer and the Smith predictor. The model-free observer estimates the current state while handling modeling errors and uncertainties of the system. The Smith predictor moves the effect of time delay outside the control closed-loop using the estimated delayed output and the actual output of the plant. The proposed method is applied to a cooperative on-ramp merging problem. First, an ultra-local model predictive control is implemented to provide a computationally efficient online speed planner agnostic to the vehicle dynamics. After that, a model-free observer is designed to estimate the current state. Finally, the proposed architecture is tested against a time-varying output delay with an upper bound of 200 milliseconds. The results demonstrate the effectiveness of the proposed method with improved tracking of intervehicle distance.

Waleed khan, Muhammad [The University of Texas at ↗

A transfer learning approach to energy-efficient control of small and medium-sized commercial buildings

Model-free reinforcement learning (RL) provides a data-driven and adaptive approach to optimize building energy use while satisfying occupant comfort. This powerful tool does not need any prior knowledge about the environment and system it is optimizing and can adapt its policy based on the changes in captures. Like any other data-driven tool, it faces high training costs due to the extensive agent-environment interactions required to capture long-term building dynamics and user comfort. Transfer learning, particularly policy distillation, offers a promising way to accelerate training by leveraging pretrained RL agents in different building and system types. Here, this study investigates online student distillation, in which the student model updates its neural network weights using outputs from teacher models. The work introduces a student distillation strategy designed for efficient knowledge transfer, along with a teacher selection method that ensures high-quality guidance. The approach is validated using a highly calibrated whole building energy model for a small/medium commercial building test facility. Results show substantial reductions in training time and data requirements while surpassing the performance of ASHRAE Guideline 36, an advanced rule-based control strategy. The distilled RL model required 45% less data and achieved 20% higher cumulative rewards than a state-of-the-art RL model, with faster convergence and lower energy consumption. These outcomes demonstrate that effective transfer learning enables a scalable and data-efficient energy management solution for commercial buildings.

ASHRAE guideline 36↗

Development of algorithms for augmenting and replacing conventional process control using reinforcement learning

Here, this work seeks to allow for the online operation and training of model-free reinforcement learning (RL) agents but limit the risk to system equipment and personnel. The parallel implementation of RL alongside more conventional process control (CPC) allows for the RL algorithm to learn from CPC. The past performance of both methods are assessed on a continuous basis allowing for a transition from CPC to RL and, if needed, transitioning back to CPC from RL. This allows for the RL algorithm to slowly and safely assume control of the process without significant degradation in control performance. It is shown that the RL can derive a near optimal policy even when coupled with a suboptimal CPC. It is also demonstrated that the coupled RL-CPC algorithm learns at a faster rate than traditional RL methods of exploration while the algorithm’s performance does not deteriorate below CPC, even when exposed to an unknown operating condition.

30 DIRECT ENERGY CONVERSION↗

Reinforcement learning pulses for transmon qubit entangling gates

The utility of a quantum computer is highly dependent on the ability to reliably perform accurate quantum logic operations. For finding optimal control solutions, it is of particular interest to explore model-free approaches, since their quality is not constrained by the limited accuracy of theoretical models for the quantum processor—in contrast to many established gate implementation strategies. In this work, we utilize a continuous control reinforcement learning algorithm to design entangling two-qubit gates for superconducting qubits; specifically, our agent constructs cross-resonance and CNOT gates without any prior information about the physical system. Using a simulated environment of fixed-frequency fixed-coupling transmon qubits, we demonstrate the capability to generate novel pulse sequences that outperform the standard cross-resonance gates in both fidelity and gate duration, while maintaining a comparable susceptibility to stochastic unitary noise. We further showcase an augmentation in training and input information that allows our agent to adapt its pulse design abilities to drifting hardware characteristics, importantly, with little to no additional optimization. Our results exhibit clearly the advantages of unbiased adaptive-feedback learning-based optimization methods for transmon gate design.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

Cooperative Merging via Online Speed Replanning: A Model-Free Approach With Vehicle-to-Vehicle Communication Packet Drop Compensation

On-ramp merging is a critical bottleneck in freeway traffic flow, contributing to congestion, accidents, and excessive fuel consumption. Although traditional ramp metering provides macroscopic control, it lacks the granularity for optimizing an individual vehicle’s trajectory. Cooperative merging, enabled by connected and automated vehicles, can potentially enhance traffic efficiency, safety, and fuel economy. However, existing research often neglects the influence of heterogeneous vehicle dynamics, unreliable vehicle-to-vehicle (V2V) communication, and real-time implementation challenges. Here, this paper introduces novel model-free online speed planners for cooperative on-ramp merging. The planners address these limitations by being agnostic to vehicle dynamics, effectively compensating for V2V communication packet drops and incurring only a light computational burden. Comprehensive evaluation, conducted on a real-time traffic-vehicle-communication co-simulation platform integrating high-fidelity vehicle dynamics, a traffic simulator, and recorded V2V communication footprints, demonstrates the effectiveness of the proposed speed planners. Simulation results reveal that the proposed method yields accurate tracking of desired speed and inter-vehicle distance, maintaining low fuel consumption even under high packet drop ratios, and demonstrating real-time implementation efficiency.

Wang, Zejiang [Univ. of Texas at Dallas, Richardso↗

Model-free stabilization via Extremum Seeking using a cost neural estimator

In this paper, a fully model-free architecture for vertical stabilization of thermonuclear plasmas in tokamak experimental reactors is presented. For the first time, an Extremum Seeking control algorithm is combined with neural networks to estimate the Lyapunov function to be minimized, resulting in a fully data-driven control architecture. The performance of different neural networks are compared. Specifically, Multilayer Perceptrons and Extreme Learning Machines are considered. The proposed architecture is tested in simulation to show that it can counteract relevant plasma disturbances, resulting in a significant improvement in terms of the achievable operative space compared to the Extremum Seeking algorithm, which still relies on model-based cost estimator.

42 ENGINEERING↗

Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators

Abstract Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The underlying problem enforces strict constraints on both individual states and actions as well as cumulative (global) constraints on energy requirements of the beam. Using historical accelerator data, we develop a physics-based surrogate model which is differentiable and allows for back-propagation of gradients. The results are evaluated in the form of a Pareto-front with two objectives. We show that the DDRL outperforms MFRL, BO, and GA on high dimensional problems.

43 PARTICLE ACCELERATORS↗

Model-free distributed learning

Model-free learning for synchronous and asynchronous quasi-static networks is presented. The network weights are continuously perturbed, while the time-varying performance index is measured and correlated with the perturbation signals; the correlation output determines the changes in the weights. The perturbation may be either via noise sources or orthogonal signals. The invariance to detailed network structure mitigates large variability between supposedly identical networks as well as implementation defects. This local, regular, and completely distributed mechanism requires no central control and involves only a few global signals. Thus it allows for integrated on-chip learning in large analog and optical networks.

Dembo, Amir↗