Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Learning and tuning fuzzy logic controllers through reinforcements

This paper presents a new method for learning and tuning a fuzzy logic controller based on reinforcements from a dynamic system. In particular, our generalized approximate reasoning-based intelligent control (GARIC) architecture (1) learns and tunes a fuzzy logic controller even when only weak reinforcement, such as a binary failure signal, is available; (2) introduces a new conjunction operator in computing the rule strengths of fuzzy control rules; (3) introduces a new localized mean of maximum (LMOM) method in combining the conclusions of several firing control rules; and (4) learns to produce real-valued control actions. Learning is achieved by integrating fuzzy inference into a feedforward neural network, which can then adaptively improve performance by using gradient descent methods. We extend the AHC algorithm of Barto et al. (1983) to include the prior control knowledge of human operators. The GARIC architecture is applied to a cart-pole balancing system and demonstrates significant improvements in terms of the speed of learning and robustness to changes in the dynamic system's parameters over previous schemes for cart-pole balancing.

Berenji, Hamid R.↗

CLEANing the Reward: Counterfactual Actions to Remove Exploratory Action Noise in Multiagent Learning

Learning in multiagent systems can be slow because agents must learn both how to behave in a complex environment and how to account for the actions of other agents. The inability of an agent to distinguish between the true environmental dynamics and those caused by the stochastic exploratory actions of other agents creates noise in each agent's reward signal. This learning noise can have unforeseen and often undesirable effects on the resultant system performance. We define such noise as exploratory action noise, demonstrate the critical impact it can have on the learning process in multiagent settings, and introduce a reward structure to effectively remove such noise from each agent's reward signal. In particular, we introduce Coordinated Learning without Exploratory Action Noise (CLEAN) rewards and empirically demonstrate their benefits

Reinforcement Learning↗

Monte Carlo Tree Search for Integrated Planning, Learning, and Execution in Nondeterministic Python

We present a novel use of Monte Carlo Tree Search (MCTS),adapted to explore a search space produced by the choice points embedded in Python code. The choice points are non-deterministic assignment statements and subroutine calls. We present MCTS extensions required for doing tree search in this context which includes control constructs like hierarchical decomposition (subroutine calls), iterative while loops and conditional statements. We demonstrate how the system works in a simulated rideshare scenario in an urban setting, and present preliminary experiments as a proof of concept.

Automatic planning↗

Towards the Development of a Multi-Agent Cognitive Networking System for the Lunar Environment

This paper details the development of a multi-agent cognitive system intended to optimize networking performance in the lunar environment. NASA’s current concept of the future of lunar communication, LunaNet [1], outlines a complex network of networks. Challenges such as scalability, interoperability and reliability must first be addressed to successfully fulfill this vision. Machine intelligence can greatly reduce the reliance on human operators and enable efficient operations for tasks such as scheduling and network management. The application of machine learning, artificial intelligence, and other automated decision-making techniques can be used to allow network nodes to intelligently sense and adapt to changes in the environment such as link disruptions, new nodes joining the network, and support for a diverse range of protocols. Cognitive networking seeks to evolve these technologies into an autonomous system with improved science data return, reliability, and scalability. In this paper, we study three main areas a means to further develop cognitive networking capabilities: networking and flight software development, analysis of wireless data for modeling and simulation, and development of algorithms for a multi-agent system.

Cognitive Networking↗

Towards the Development of a Multi-Agent Cognitive Networking System for the Lunar Environment

This paper details the development of a multi-agent cognitive system intended to optimize networking performance in the lunar environment. One concept of the future of lunar communication, LunaNet, outlines a complex network of networks. Challenges such as scalability, interoperability, and reliability must first be addressed to successfully fulfill this vision. Machine intelligence can greatly reduce the reliance on human operators and enable efficient operations for tasks such as scheduling and network management. Machine learning, artificial intelligence, and other automated decision-making techniques can be used to allow network nodes to intelligently sense and adapt to changes in the environment such as link disruptions, new nodes joining the network, and support for a diverse range of protocols. Cognitive networking seeks to evolve these technologies into an autonomous system with improved science data return, reliability, and scalability. In this paper, we study four main areas as a means to further develop cognitive networking capabilities: networking protocol development, analysis of wireless data for modeling and simulation, development of algorithms for a multi-agent system, and spectrum sensing technology.

cognitive networking↗

Minority University System Engineering: A Small Satellite Design Experience Held at the Jet Propulsion Laboratory During the Summer of 1996

The University of Texas at El Paso (UTEP) in conjunction with the Jet Propulsion Laboratory (JPL), North Carolina A&T and California State University of Los Angeles participated during the summer of 1996 in a prototype program known as Minority University Systems Engineering (MUSE). The program consisted of a ten week internship at JPL for students and professors of the three universities. The purpose of MUSE as set forth in the MUSE program review August 5, 1996 was for the participants to gain experience in the following areas: 1) Gain experience in a multi-disciplinary project; 2) Gain experience working in a culturally diverse atmosphere; 3) Provide field experience for students to reinforce book learning; and 4) Streamline the design process in two areas: make it more financially feasible; and make it faster.

Ordaz, Miguel Angel↗

Toward applied behavior analysis of life aloft

This article deals with systems at multiple levels, at least from cell to organization. It also deals with learning, decision making, and other behavior at multiple levels. Technological development of a human behavioral ecosystem appropriate to space environments requires an analytic and synthetic orientation, explicitly experimental in nature, dictated by scientific and pragmatic considerations, and closely approximating procedures of established effectiveness in other areas of natural science. The conceptual basis of such an approach has its roots in environmentalism which has two main features: (1) knowledge comes from experience rather than from innate ideas, divine revelation, or other obscure sources; and (2) action is governed by consequences rather than by instinct, reason, will, beliefs, attitudes or even the currently fashionable cognitions. Without an experimentally derived data base founded upon such a functional analysis of human behavior, the overgenerality of "ecological systems" approaches render them incapable of ensuring the successful establishment of enduring space habitats. Without an experimentally derived function account of individual behavioral variability, a natural science of behavior cannot exist. And without a natural science of behavior, the social sciences will necessarily remain in their current status as disciplines of less than optimal precision or utility. Such a functional analysis of human performance should provide an operational account of behavior change in a manner similar to the way in which Darwin's approach to natural selection accounted for the evolution of phylogenetic lines (i.e., in descriptive, nonteleological terms). Similarly, as Darwin's account has subsequently been shown to be consonant with information obtained at the cellular level, so too should behavior principles ultimately prove to be in accord with an account of ontogenetic adaptation at a biochemical level. It would thus seem obvious that the most productive conceptual and methodological approaches to long-term research investments focused upon human behavior in space environments will require multidisciplinary inputs from such wide-ranging fields as molecular biology, environmental physiology, behavioral biology, architecture, sociology, and political science, among others.

Review↗

UAS Conflict-Avoidance Using Multiagent RL with Abstract Strategy Type Communication

The use of unmanned aerial systems (UAS) in the national airspace is of growing interest to the research community. Safety and scalability of control algorithms are key to the successful integration of autonomous system into a human-populated airspace. In order to ensure safety while still maintaining efficient paths of travel, these algorithms must also accommodate heterogeneity of path strategies of its neighbors. We show that, using multiagent RL, we can improve the speed with which conflicts are resolved in cases with up to 80 aircraft within a section of the airspace. In addition, we show that the introduction of abstract agent strategy types to partition the state space is helpful in resolving conflicts, particularly in high congestion.

Unmanned Autonomous Systems↗

Announced Strategy Types in Multiagent RL for Conflict-Avoidance in the National Airspace

The use of unmanned aerial systems (UAS) in the national airspace is of growing interest to the research community. Safety and scalability of control algorithms are key to the successful integration of autonomous system into a human-populated airspace. In order to ensure safety while still maintaining efficient paths of travel, these algorithms must also accommodate heterogeneity of path strategies of its neighbors. We show that, using multiagent RL, we can improve the speed with which conflicts are resolved in cases with up to 80 aircraft within a section of the airspace. In addition, we show that the introduction of abstract agent strategy types to partition the state space is helpful in resolving conflicts, particularly in high congestion.

National Airspace↗

AdaStress

This is a tutorial on AdaStress, a tool for finding and analyzing the likeliest failures in a simulated system under test. The presentation outlines the adaptive stress testing framework, provides a demonstration of use, and showcases several examples of failure detection in a complex real-world system.

Reinforcement learning↗

Transient Optimization for the Betterment of Turbine Electrified Energy Management

Gas turbine engine transients are associated with degraded compressor operability, which must be addressed by the engine control system and accounted for in the engine design. Failure to do so may result in events such as compressor stall/surge and combustor blow out. Transient operability concerns constrain the engine design and can result in sacrifices of efficiency and/or thrust responsiveness. The traditional approach to transient operability management is control logic that limits the fuel flow command. A companion paper presents a strategy for optimizing the transient fuel flow control logic taking into consideration transient operability and thrust responsiveness. The study covered here extends this idea to an electrified gas turbine engine that employs a power/energy management concept known as Turbine Electrified Energy Management (TEEM). TEEM uses an electric power system interfaced with the engine (hence the term ‘electrified gas turbine engine’) to further improve transient operability and alleviate associated design constraints. There can be costs associated with implementing TEEM in terms of power and energy requirements that impact the size of the electrical power system. However, the results of this study show that through optimization of the transient limit logic, power and energy requirements needed to implement TEEM can be significantly reduced. Among the conclusions that can be drawn from the results of the illustrative application covered herein are: (1) there is a reduction in the electric machine power requirement to manage operability during accelerations by 200 to 400 hp, and (2) power transfer from the low pressure spool (LPS) to the high pressure spool (HPS) is the most effective option for improving operability during decelerations, followed by the options of only injecting power on the HPS or only extracting power from the LPS.

Turbine Electrified Energy Management↗

Transient Optimization for the Betterment of Turbine Electrified Energy Management

Gas turbine engine transients are associated with degraded compressor operability, which must be addressed by the engine control system and accounted for in the engine design. Failure to do so may result in events such as compressor stall/surge and combustor blow out. Transient operability concerns constrain the engine design and can result in sacrifices of efficiency and/or thrust responsiveness. The traditional approach to transient operability management is control logic that limits the fuel flow command. A companion paper presents a strategy for optimizing the transient fuel flow control logic taking into consideration transient operability and thrust responsiveness. The study covered here extends this idea to an electrified gas turbine engine that employs a power/energy management concept known as Turbine Electrified Energy Management (TEEM). TEEM uses an electric power system interfaced with the engine (hence the term ‘electrified gas turbine engine’) to further improve transient operability and alleviate associated design constraints. There can be costs associated with implementing TEEM in terms of power and energy requirements that impact the size of the electrical power system. However, the results of this study show that through optimization of the transient limit logic, power and energy requirements needed to implement TEEM can be significantly reduced. Among the conclusions that can be drawn from the results of the illustrative application covered herein are: (1) there is a reduction in the electric machine power requirement to manage operability during accelerations by 200 to 400 hp, and (2) power transfer from the low pressure spool (LPS) to the high pressure spool (HPS) is the most effective option for improving operability during decelerations, followed by the options of only injecting power on the HPS or only extracting power from the LPS.

transient↗