Search NASA⌕ Search

Engineering topics

Char, Ian

Publications and source records attributed to Char, Ian.

Automated experimental design of safe rampdowns via probabilistic machine learning

Abstract Typically the rampdown phase of a shot consists of a decrease in current and injected power and optionally a change in shape, but there is considerable flexibility in the rate, sequencing, and duration of these changes. On the next generation of tokamaks it is essential that this is done safely as the device could be damaged by the stored thermal and electromagnetic energy present in the plasma. This works presents a procedure for automatically choosing experimental rampdown designs to rapidly converge to an effective rampdown trajectory. This procedure uses probabilistic machine learning methods paired with acquisition functions taken from Bayesian optimization. In a set of 2022 experiments at DIII-D, the rampdown designs produced by our method maintained plasma control down to substantially lower current and energy levels than are typically observed. The actions predicted by the model significantly improved as the model was able to explore over the course of the experimental campaign.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making

One of the great challenges with decision making tasks on real world systems is the fact that data is sparse and acquiring additional data is expensive. In these cases, it is often crucial to make a model of the environment to assist in making decisions. At the same time, limited data means that learned models are erroneous, making it just as important to equip the model with good predictive uncertainties. In the context of learning sequential decision making policies, these uncertainties can prove useful for informing which data to collect for the greatest improvement in policy performance \citep{mehta2021experimental, mehta2022exploration} or informing the policy about unsure regions of state and action space to avoid during test time \citep{yu2020mopo}. Additionally, assuming that realistic samples of the environment can be drawn, an adaptable policy can be trained that attempts to make optimal decisions for any given possible instance of the environment \citep{ghosh2022offline, chen2021offline}. In this work, we examine the so-called ``probabilistic neural network'' (PNN) model that is ubiquitous in model-based reinforcement learning (MBRL) works. We argue that while PNN models may have good marginal uncertainties, they form a distribution of non-smooth transition functions. Not only are these samples unrealistic and may hamper adaptability, but we also assert that this leads to poor uncertainty estimates when predicting multiple step trajectory estimates. To address this issue, we propose a simple sampling method that can be implemented on top of pre-existing models.We evaluate our sampling technique on a number of environments, including a realistic nuclear fusion task, and find that, not only do smooth transition function samples produce more calibrated uncertainties, but they also lead to better downstream performance for an adaptive policy.

Offline Reinforcement Learning↗