Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Transient Optimization for the Betterment of Turbine Electrified Energy Management

Gas turbine engine transients are associated with degraded compressor operability, which must be addressed by the engine control system and accounted for in the engine design. Failure to do so may result in events such as compressor stall/surge and combustor blow out. Transient operability concerns constrain the engine design and can result in sacrifices of efficiency and/or thrust responsiveness. The traditional approach to transient operability management is control logic that limits the fuel flow command. A companion paper presents a strategy for optimizing the transient fuel flow control logic taking into consideration transient operability and thrust responsiveness. The study covered here extends this idea to an electrified gas turbine engine that employs a power/energy management concept known as Turbine Electrified Energy Management (TEEM). TEEM uses an electric power system interfaced with the engine (hence the term ‘electrified gas turbine engine’) to further improve transient operability and alleviate associated design constraints. There can be costs associated with implementing TEEM in terms of power and energy requirements that impact the size of the electrical power system. However, the results of this study show that through optimization of the transient limit logic, power and energy requirements needed to implement TEEM can be significantly reduced. Among the conclusions that can be drawn from the results of the illustrative application covered herein are: (1) there is a reduction in the electric machine power requirement to manage operability during accelerations by 200 to 400 hp, and (2) power transfer from the low pressure spool (LPS) to the high pressure spool (HPS) is the most effective option for improving operability during decelerations, followed by the options of only injecting power on the HPS or only extracting power from the LPS.

transient↗

Route-Recapturing State-Based Horizontal Maneuver Strategy for Automated Detect-and-Avoid

This report describes a novel approach to the development of a horizontal maneuver guidance strategy for Detect-and-Avoid systems. The maneuver guidance strategy provides a directive turn action that can be automatically executed by the vehicle’s auto-pilot system, taking into account the cost of recapturing the flight plan path. Pairwise conflict scenarios with non-accelerating intruders are simulated to validate the effectiveness of the maneuver guidance strategy. Initial results suggest the strategy is more effective for faster ownship than for slower ownship, which is unable to avoid conflict in certain scenarios against fast intruders. These findings indicate this novel approach shows great potential, but improvement to its performance is necessary and will be future work.

detect-and-avoid↗

Evaluating a Cognitive Extension for the Licklider Transmission Protocol in a Spacecraft Emulation Testbed

In space communications, particularly when involving regions beyond cislunar space, the development of advanced networking solutions is essential to address the challenges posed by limited connectivity, substantial propagation delays, and radio signal variations. This study explores a data-driven approach to the Licklider Transmission Protocol (LTP), specifically focusing on dynamically adjusting the maximum payload size of segments. Prior research has emphasized the potential benefits of dynamically adjusting this parameter, introducing the concept of Cognitive LTP. This paper presents a software implementation of Cognitive LTP (CLTP) within an open-source Delay Tolerant Networking (DTN) framework, specifically the High-rate Delay Tolerant Networking (HDTN), and experimentally evaluates its performance under realistic space conditions. Leveraging the Cognitive Ground Testbed (CGT), developed by NASA GRC for spacecraft communication emulation, this study effectively bridges the gap between theoretical advancements and practical applications. By thoroughly analyzing CLTP’s functionality within the CGT, this research offers insights into the practical implications of adaptive networking strategies, emphasizing the importance of conducting tests in relevant environments for the maturation of space communication technologies.

Delay Tolerant Networking↗

Fuzzy self-learning control for magnetic servo system

It is known that an effective control system is the key condition for successful implementation of high-performance magnetic servo systems. Major issues to design such control systems are nonlinearity; unmodeled dynamics, such as secondary effects for copper resistance, stray fields, and saturation; and that disturbance rejection for the load effect reacts directly on the servo system without transmission elements. One typical approach to design control systems under these conditions is a special type of nonlinear feedback called gain scheduling. It accommodates linear regulators whose parameters are changed as a function of operating conditions in a preprogrammed way. In this paper, an on-line learning fuzzy control strategy is proposed. To inherit the wealth of linear control design, the relations between linear feedback and fuzzy logic controllers have been established. The exercise of engineering axioms of linear control design is thus transformed into tuning of appropriate fuzzy parameters. Furthermore, fuzzy logic control brings the domain of candidate control laws from linear into nonlinear, and brings new prospects into design of the local controllers. On the other hand, a self-learning scheme is utilized to automatically tune the fuzzy rule base. It is based on network learning infrastructure; statistical approximation to assign credit; animal learning method to update the reinforcement map with a fast learning rate; and temporal difference predictive scheme to optimize the control laws. Different from supervised and statistical unsupervised learning schemes, the proposed method learns on-line from past experience and information from the process and forms a rule base of an FLC system from randomly assigned initial control rules.

Tarn, J. H.↗

Emergence of relations and the essence of learning: a review of Sidman's Equivalence relations and behavior: a research story. Book review

The author reviews and comments on the book Equivalence relations and behavior: a research story by Murray Sidman. Sidman's book reports his research about equivalence relations and competencies in children with mental retardation and how it relates to behavior. Sidman used the idea of stimulus-stimulus relations among features of the environment to develop his theories about equivalence relations. Experimental work with children and animals demonstrated their ability to use equivalence relations to learn new tasks. The subject received feedback and reinforcement for specific choices made during training, then was presented with new choices during testing. Results of the tests indicate that subjects were able to establish relations and retrieve them in different situations.

NASA Discipline Space Human Factors↗

An Onboard ISS Virtual Reality Trainer

Prior to the retirement of the Space Shuttle, many exterior repairs on the International Space Station (ISS) were carried out by shuttle astronauts, trained on the ground and flown to the Station to perform these specific repairs. With the retirement of the shuttle, this is no longer an available option. As such, the need for ISS crew members to review scenarios while on flight, either for tasks they already trained for on the ground or for contingency operations has become a very critical issue. NASA astronauts prepare for Extra-Vehicular Activities (EVA) or Spacewalks through numerous training media, such as: self-study, part task training, underwater training in the Neutral Buoyancy Laboratory (NBL), hands-on hardware reviews and training at the Virtual Reality Laboratory (VRLab). In many situations, the time between the last session of a training and an EVA task might be 6 to 8 months. EVA tasks are critical for a mission and as time passes the crew members may lose proficiency on previously trained tasks and their options to refresh or learn a new skill while on flight are limited to reading training materials and watching videos. In addition, there is an increased need for unplanned contingency repairs to fix problems arising as the Station ages. In order to help the ISS crew members maintain EVA proficiency or train for contingency repairs during their mission, the Johnson Space Center's VRLab designed an immersive ISS Virtual Reality Trainer (VRT). The VRT incorporates a unique optical system that makes use of the already successful Dynamic On-board Ubiquitous Graphics (DOUG) software to assist crew members with procedure reviews and contingency EVAs while on board the Station. The need to train and re-train crew members for EVAs and contingency scenarios is crucial and extremely demanding. ISS crew members are now asked to perform EVA tasks for which they have not been trained and potentially have never seen before. The Virtual Reality Trainer (VRT) provides an immersive 3D environment similar to the one experienced at the VRLab crew training facility at the NASA Johnson Space Center. VRT bridges the gap by allowing crew members to experience an interactive, 3D environment to reinforce skills already learned and to explore new work sites and repair procedures outside the Station.

Miralles, Evelyn↗

Evolving fuzzy rules in a learning classifier system

The fuzzy classifier system (FCS) combines the ideas of fuzzy logic controllers (FLC's) and learning classifier systems (LCS's). It brings together the expressive powers of fuzzy logic as it has been applied in fuzzy controllers to express relations between continuous variables, and the ability of LCS's to evolve co-adapted sets of rules. The goal of the FCS is to develop a rule-based system capable of learning in a reinforcement regime, and that can potentially be used for process control.

Valenzuela-Rendon, Manuel↗

Attitude determination using an adaptive multiple model filtering Scheme

Attitude determination has been considered as a permanent topic of active research and perhaps remaining as a forever-lasting interest for spacecraft system designers. Its role is to provide a reference for controls such as pointing the directional antennas or solar panels, stabilizing the spacecraft or maneuvering the spacecraft to a new orbit. Least Square Estimation (LSE) technique was utilized to provide attitude determination for the Nimbus 6 and G. Despite its poor performance (estimation accuracy consideration), LSE was considered as an effective and practical approach to meet the urgent need and requirement back in the 70's. One reason for this poor performance associated with the LSE scheme is the lack of dynamic filtering or 'compensation'. In other words, the scheme is based totally on the measurements and no attempts were made to model the dynamic equations of motion of the spacecraft. We propose an adaptive filtering approach which employs a bank of Kalman filters to perform robust attitude estimation. The proposed approach, whose architecture is depicted, is essentially based on the latest proof on the interactive multiple model design framework to handle the unknown of the system noise characteristics or statistics. The concept fundamentally employs a bank of Kalman filter or submodel, instead of using fixed values for the system noise statistics for each submodel (per operating condition) as the traditional multiple model approach does, we use an on-line dynamic system noise identifier to 'identify' the system noise level (statistics) and update the filter noise statistics using 'live' information from the sensor model. The advanced noise identifier, whose architecture is also shown, is implemented using an advanced system identifier. To insure the robust performance for the proposed advanced system identifier, it is also further reinforced by a learning system which is implemented (in the outer loop) using neural networks to identify other unknown quantities such as spacecraft dynamics parameters, gyro biases, dynamic disturbances, or environment variations.

Lam, Quang↗

Initial Approach to Collect Small Unmanned Aircraft System Off-Nominal Operational Situations Data

NASA is developing the Unmanned Aircraft System Traffic Management research platform to safely integrate small unmanned aircraft operations in large-scale at low-altitudes. As a part of this effort, small unmanned aircraft system off-nominal operational situations data collection process has been developed to take lessons learned and to reinforce operational compliance. In this paper, descriptions of variables used for digital data collection and an online report form for collection of observational data from the operators (contextual data) are provided. They are used to collect off-nominal data from the Unmanned Aircraft System Traffic Management National Campaign in 2017. The digital data show that 2 out of 118 campaign operations (1.7%) encountered loss of navigation. Since the campaign aircraft used Global Positioning System for navigation, it is likely that unobstructed view of the sky at the campaign locations contributed to this small number. Also, 4 out of 47 operations (8.5%) encountered loss of communications. A relatively short distance between ground control system and aircraft, ranging from 2300 feet to 4200 feet, likely contributed to this small number. There was no data to identify the loss of communications condition, aircraft received signal strength, for the remaining 71 operations suggesting that some operators may not be monitoring unmanned aircraft communications system performance or monitoring it with different parameters. For the contextual data, due to the low number of total reports during the campaign, no significant trends emerged. This is an initial attempt to collect contextual data from small unmanned aircraft operators about off-nominal situations, and changes will be made to the future data collection to improve the amount and quality of the information.

unmanned aviation systems traffic management (UTM)↗

Initial Approach to Collect Small Unmanned Aircraft System Off-Nominal Operational Situations Data

NASA is developing the Unmanned Aircraft System Traffic Management research platform to safely integrate small unmanned aircraft operations in large-scale at low-altitudes. As a part of this effort, small unmanned aircraft system off-nominal operational situations data collection process has been developed to take lessons learned and to reinforce operational compliance. In this paper, descriptions of variables used for digital data collection and an online report form for collection of observational data from the operators (contextual data) are provided. They are used to collect off-nominal data from the Unmanned Aircraft System Traffic Management National Campaign in 2017. The digital data show that 2 out of 118 campaign operations (1.7%) encountered loss of navigation. Since the campaign aircraft used Global Positioning System for navigation, it is likely that unobstructed view of the sky at the campaign locations contributed to this small number. Also, 4 out of 47 operations (8.5%) encountered loss of communications. A relatively short distance between ground control system and aircraft, ranging from 2300 feet to 4200 feet, likely contributed to this small number. There was no data to identify the loss of communications condition, aircraft received signal strength, for the remaining 71 operations suggesting that some operators may not be monitoring unmanned aircraft communications system performance or monitoring it with different parameters. For the contextual data, due to the low number of total reports during the campaign, no significant trends emerged. This is an initial attempt to collect contextual data from small unmanned aircraft operators about off-nominal situations, and changes will be made to the future data collection to improve the amount and quality of the information.

Jung, Jaewoo↗

Periodic shock with added clock.

Periodic shock in rats maintained on variable-interval schedule of reinforcement interspersed with sequence of three different stimulus conditions

CONDITIONING (LEARNING)↗

Using rewards and penalties to obtain desired subject performance

Operant conditioning procedures, specifically the use of negative reinforcement, in achieving stable learning behavior is described. The critical tracking test (CTT) a method of detecting human operator impairment was tested. A pass level is set for each subject, based on that subject's asymptotic skill level while sober. It is critical that complete training take place before the individualized pass level is set in order that the impairment can be detected. The results provide a more general basis for the application of reward/penalty structures in manual control research.

Cook, M.↗

Pavlovian, Skinner, and Other Behaviourists' Contributions to AI

A version of the definition of intelligent behaviour will be supplied in the context of real and artificial systems. Short presentation of principles of learning, starting with Pavlovian s classical conditioning through reinforced response and operant conditioning of Thorndike and Skinner and finishing with cognitive learning of Tolman and Bandura will be given. The most important figures within behaviourism, especially those with contribution to AI, will be described. Some tools of artificial intelligence that act according to those principles will be presented. An attempt will be made to show when some simple rules for behaviour modifications can lead to a complex intelligent behaviour.

Kosinski, Withold↗

Rapid motor learning in the translational vestibulo-ocular reflex

Motor learning was induced in the translational vestibulo-ocular reflex (TVOR) when monkeys were repeatedly subjected to a brief (0.5 sec) head translation while they tried to maintain binocular fixation on a visual target for juice rewards. If the target was world-fixed, the initial eye speed of the TVOR gradually increased; if the target was head-fixed, the initial eye speed of the TVOR gradually decreased. The rate of learning acquisition was very rapid, with a time constant of approximately 100 trials, which was equivalent to <1 min of accumulated stimulation. These learned changes were consolidated over >or=1 d without any reinforcement, indicating induction of long-term synaptic plasticity. Although the learning generalized to targets with different viewing distances and to head translations with different accelerations, it was highly specific for the particular combination of head motion and evoked eye movement associated with the training. For example, it was specific to the modality of the stimulus (translation vs rotation) and the direction of the evoked eye movement in the training. Furthermore, when one eye was aligned with the heading direction so that it remained motionless during training, learning was not expressed in this eye, but only in the other nonaligned eye. These specificities show that the learning sites are neither in the sensory nor the motor limb of the reflex but in the sensory-motor transformation stage of the reflex. The dependence of the learning on both head motion and evoked eye movement suggests that Hebbian learning may be one of the underlying cellular mechanisms.

Non-NASA Center↗

Program Helps Simulate Neural Networks

Neural Network Environment on Transputer System (NNETS) computer program provides users high degree of flexibility in creating and manipulating wide variety of neural-network topologies at processing speeds not found in conventional computing environments. Supports back-propagation and back-propagation-related algorithms. Back-propagation algorithm used is implementation of Rumelhart's generalized delta rule. NNETS developed on INMOS Transputer(R). Predefines back-propagation network, Jordan network, and reinforcement network to assist users in learning and defining own networks. Also enables users to configure other neural-network paradigms from NNETS basic architecture. Small portion of software written in OCCAM(R) language.

Villarreal, James↗

Assessing the Effects of Momentary Priming on Memory Retention During an Interference Task

A memory aid, that used brief (33ms) presentations of previously learned information (target words), was assessed on its ability to reinforce memory for target words while the subject was performing an interference task. The interference task required subjects to learn new words and thus interfered with their memory of the target words. The brief presentation (momentary memory priming) was hypothesized to refresh the subjects memory of the target words. 143 subjects, in a within subject design, were given a 33ms presentation of the target memory words during the interference task in a treatment condition and a blank 33ms presentation in the control condition. The primary dependent measure, memory loss over the interference trial, was not significantly different between the two conditions. The memory prime did not appear to hinder the subjects performance on the interference task. This paper describes the experiment and the results along with suggestions for future research.

Schutte, Paul C.↗