Search NASA⌕ Search

SEARCH · Search NASA

Results for “Inverse Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Use of Inverse Reinforcement Learning for Identity Prediction

We adopt Markov Decision Processes (MDP) to model sequential decision problems, which have the characteristic that the current decision made by a human decision maker has an uncertain impact on future opportunity. We hypothesize that the individuality of decision makers can be modeled as differences in the reward function under a common MDP model. A machine learning technique, Inverse Reinforcement Learning (IRL), was used to learn an individual's reward function based on limited observation of his or her decision choices. This work serves as an initial investigation for using IRL to analyze decision making, conducted through a human experiment in a cyber shopping environment. Specifically, the ability to determine the demographic identity of users is conducted through prediction analysis and supervised learning. The results show that IRL can be used to correctly identify participants, at a rate of 68% for gender and 66% for one of three college major categories.

Hayes, Roy↗

Ground Delay Program Analytics with Behavioral Cloning and Inverse Reinforcement Learning

We used historical data to build two types of model that predict Ground Delay Program implementation decisions and also produce insights into how and why those decisions are made. More specifically, we built behavioral cloning and inverse reinforcement learning models that predict hourly Ground Delay Program implementation at Newark Liberty International and San Francisco International airports. Data available to the models include actual and scheduled air traffic metrics and observed and forecasted weather conditions. We found that the random forest behavioral cloning models we developed are substantially better at predicting hourly Ground Delay Program implementation for these airports than the inverse reinforcement learning models we developed. However, all of the models struggle to predict the initialization and cancellation of Ground Delay Programs. We also investigated the structure of the models in order to gain insights into Ground Delay Program implementation decision making. Notably, characteristics of both types of model suggest that GDP implementation decisions are more tactical than strategic: they are made primarily based on conditions now or conditions anticipated in only the next couple of hours.

Bloem, Michael↗

Toward Justifiable Trust in Autonomous Systems Incorporating Human Knowledge in Autonomous Systems through Machine Learning

Trust in Autonomous Systems is largely about humans trusting the decisions made by autonomous systems. This trust can be increased through learning from domain experts. In particular, autonomous systems can learn offline from past mission operations before conducting any operations of its own. Additionally, autonomous systems can learn online by obtaining human feedback during operations. We will discuss several classes of machine learning methods and our application of them to autonomous systems. The first class of methods is anomaly detection, which uses operations data to identify examples of anomalous operations. The second class of methods is inverse reinforcement learning, also known as apprenticeship learning, that takes past operations data as input and yields a controller that is able to duplicate the operations described by the data. The third class is active learning, which identifies examples on which the model is most uncertain and requests domain expert feedback.

Oza, Nikunj C.↗

Towards Intelligent Control for Next Generation Aircraft

NASA Aeronautics Subsonic Fixed Wing Project is focused on mitigating the environmental and operation impacts expected as aviation operations triple by 2025. The approach is to extend technological capabilities and explore novel civil transport configurations that reduce noise, emissions, fuel consumption and field length. Two Next Generation (NextGen) aircraft have been identified to meet the Subsonic Fixed Wing Project goals - these are the Hybrid Wing-Body (HWB) and Cruise Efficient Short Take-Off and Landing (CESTOL) aircraft. The technologies and concepts developed for these aircraft complicate the vehicle s design and operation. In this paper, flight control challenges for NextGen aircraft are described. The objective of this paper is to examine the potential of state-of-the-art control architectures and algorithms to meet the challenges and needed performance metrics for NextGen flight control. A broad range of conventional and intelligent control approaches are considered, including dynamic inversion control, integrated flight-propulsion control, control allocation, adaptive dynamic inversion control, data-based predictive control and reinforcement learning control.

Acosta, Diana Michelle↗

Harnessing Collaborative Learning Automata to Guide Multi-objective Optimization based Inverse Analysis for Structural Damage Identification

Structural damage identification based on physical models is often transformed into an optimization problem that minimizes the difference between measurement information of structure being monitored and the model prediction in the parametric space. However, the objective function in this context often exhibits multimodality, involving high-dimensional variables due to the reliance on finite element models for damage identification. These features pose challenges to optimization algorithms, where entrapment in local solutions can lead to false positives and false negatives in damage identification. In this research, we propose a reinforcement learning based multi-swarm optimizer to tackle such challenges in pursuit of a small yet diverse solution set that can capture the true damage scenario as one of the solutions. The proposed method leverages the flexibility of the particle swarm optimizer and incorporates novel strategies of metaheuristics to realize targeted improvement. To enable the particle swarm to adaptively select the appropriate search strategy based on the current environment, we adopt the learning automata technique, which sidesteps the need for reward strategy selection that is usually ad hoc at each step of the search. The integration harnesses the automatic learning and self-adaptation capabilities of learning automata, enabling the particles to navigate based on environmental signals. This leads to accumulated probabilities tied to advantageous movements, fostering an adaptive exploration of particles in the search space. The proposed approach is first validated through implementing into benchmark test cases with comparisons. It is then applied to structural damage identification with piezoelectric admittance experimental signals. `The results highlight the capability of the algorithm to identify a small solution set with high accuracy to match the actual damage scenario.

Yang Zhang↗

Piezoelectric impedance-based high-accuracy damage identification using sparsity conscious multi-objective optimization inverse analysis

Two elements are essential in structural health monitoring utilizing dynamic responses: response measurement with high-frequency contents, i.e., small characteristic wavelengths, that can adequately reflect damage features, and effective inverse identification analysis that is however oftentimes under-determined. The advancement of smart structure integration has led to active interrogation through frequency-sweeping piezoelectric impedance measurement at high frequency range. In this research we develop a multi-objective optimization formulation for the identification of damage location and severity utilizing piezoelectric impedance. While one optimization objective is to match the response measurement with finite element model prediction in the damage parametric space, the other is the number of locations of damage, i.e., the sparsity of damage index as the solution vector, since damage usually occurs within a small number of locations. This multi-objective formulation fits well the under-determined nature of damage identification, as it naturally provides multiple solutions as basis for further elucidation. The challenge remaining is how to find a small solution set that can include the actual damage scenario. Here we develop a novel inverse identification framework utilizing the intelligent swarm optimizer which possesses flexibility for enhancement. We first embed a sparsity enforcement process into the population generation of the optimizer, which yields a solution repository intrinsically possessing sparsity. We then apply reinforcement learning so the agents can adaptively opt for local strategies with the aim of enriching the searching patterns to diversify the solutions. Through the incorporation of a Q-table, searching toward more promising directions will be rewarded. Our case analyses employing experimental data indicate that this sparsity-conscious multi-objective particle swarm optimization technique can lead to a small solution set which generally encompasses the true damage scenario. This effectively solves the structural damage identification problem with piezoelectric impedance measurement.

Yang Zhang↗