Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Fuzzy self-learning control for magnetic servo system

It is known that an effective control system is the key condition for successful implementation of high-performance magnetic servo systems. Major issues to design such control systems are nonlinearity; unmodeled dynamics, such as secondary effects for copper resistance, stray fields, and saturation; and that disturbance rejection for the load effect reacts directly on the servo system without transmission elements. One typical approach to design control systems under these conditions is a special type of nonlinear feedback called gain scheduling. It accommodates linear regulators whose parameters are changed as a function of operating conditions in a preprogrammed way. In this paper, an on-line learning fuzzy control strategy is proposed. To inherit the wealth of linear control design, the relations between linear feedback and fuzzy logic controllers have been established. The exercise of engineering axioms of linear control design is thus transformed into tuning of appropriate fuzzy parameters. Furthermore, fuzzy logic control brings the domain of candidate control laws from linear into nonlinear, and brings new prospects into design of the local controllers. On the other hand, a self-learning scheme is utilized to automatically tune the fuzzy rule base. It is based on network learning infrastructure; statistical approximation to assign credit; animal learning method to update the reinforcement map with a fast learning rate; and temporal difference predictive scheme to optimize the control laws. Different from supervised and statistical unsupervised learning schemes, the proposed method learns on-line from past experience and information from the process and forms a rule base of an FLC system from randomly assigned initial control rules.

Tarn, J. H.↗

Emergence of relations and the essence of learning: a review of Sidman's Equivalence relations and behavior: a research story. Book review

The author reviews and comments on the book Equivalence relations and behavior: a research story by Murray Sidman. Sidman's book reports his research about equivalence relations and competencies in children with mental retardation and how it relates to behavior. Sidman used the idea of stimulus-stimulus relations among features of the environment to develop his theories about equivalence relations. Experimental work with children and animals demonstrated their ability to use equivalence relations to learn new tasks. The subject received feedback and reinforcement for specific choices made during training, then was presented with new choices during testing. Results of the tests indicate that subjects were able to establish relations and retrieve them in different situations.

NASA Discipline Space Human Factors↗

An Onboard ISS Virtual Reality Trainer

Prior to the retirement of the Space Shuttle, many exterior repairs on the International Space Station (ISS) were carried out by shuttle astronauts, trained on the ground and flown to the Station to perform these specific repairs. With the retirement of the shuttle, this is no longer an available option. As such, the need for ISS crew members to review scenarios while on flight, either for tasks they already trained for on the ground or for contingency operations has become a very critical issue. NASA astronauts prepare for Extra-Vehicular Activities (EVA) or Spacewalks through numerous training media, such as: self-study, part task training, underwater training in the Neutral Buoyancy Laboratory (NBL), hands-on hardware reviews and training at the Virtual Reality Laboratory (VRLab). In many situations, the time between the last session of a training and an EVA task might be 6 to 8 months. EVA tasks are critical for a mission and as time passes the crew members may lose proficiency on previously trained tasks and their options to refresh or learn a new skill while on flight are limited to reading training materials and watching videos. In addition, there is an increased need for unplanned contingency repairs to fix problems arising as the Station ages. In order to help the ISS crew members maintain EVA proficiency or train for contingency repairs during their mission, the Johnson Space Center's VRLab designed an immersive ISS Virtual Reality Trainer (VRT). The VRT incorporates a unique optical system that makes use of the already successful Dynamic On-board Ubiquitous Graphics (DOUG) software to assist crew members with procedure reviews and contingency EVAs while on board the Station. The need to train and re-train crew members for EVAs and contingency scenarios is crucial and extremely demanding. ISS crew members are now asked to perform EVA tasks for which they have not been trained and potentially have never seen before. The Virtual Reality Trainer (VRT) provides an immersive 3D environment similar to the one experienced at the VRLab crew training facility at the NASA Johnson Space Center. VRT bridges the gap by allowing crew members to experience an interactive, 3D environment to reinforce skills already learned and to explore new work sites and repair procedures outside the Station.

Miralles, Evelyn↗

Evolving fuzzy rules in a learning classifier system

The fuzzy classifier system (FCS) combines the ideas of fuzzy logic controllers (FLC's) and learning classifier systems (LCS's). It brings together the expressive powers of fuzzy logic as it has been applied in fuzzy controllers to express relations between continuous variables, and the ability of LCS's to evolve co-adapted sets of rules. The goal of the FCS is to develop a rule-based system capable of learning in a reinforcement regime, and that can potentially be used for process control.

Valenzuela-Rendon, Manuel↗

Attitude determination using an adaptive multiple model filtering Scheme

Attitude determination has been considered as a permanent topic of active research and perhaps remaining as a forever-lasting interest for spacecraft system designers. Its role is to provide a reference for controls such as pointing the directional antennas or solar panels, stabilizing the spacecraft or maneuvering the spacecraft to a new orbit. Least Square Estimation (LSE) technique was utilized to provide attitude determination for the Nimbus 6 and G. Despite its poor performance (estimation accuracy consideration), LSE was considered as an effective and practical approach to meet the urgent need and requirement back in the 70's. One reason for this poor performance associated with the LSE scheme is the lack of dynamic filtering or 'compensation'. In other words, the scheme is based totally on the measurements and no attempts were made to model the dynamic equations of motion of the spacecraft. We propose an adaptive filtering approach which employs a bank of Kalman filters to perform robust attitude estimation. The proposed approach, whose architecture is depicted, is essentially based on the latest proof on the interactive multiple model design framework to handle the unknown of the system noise characteristics or statistics. The concept fundamentally employs a bank of Kalman filter or submodel, instead of using fixed values for the system noise statistics for each submodel (per operating condition) as the traditional multiple model approach does, we use an on-line dynamic system noise identifier to 'identify' the system noise level (statistics) and update the filter noise statistics using 'live' information from the sensor model. The advanced noise identifier, whose architecture is also shown, is implemented using an advanced system identifier. To insure the robust performance for the proposed advanced system identifier, it is also further reinforced by a learning system which is implemented (in the outer loop) using neural networks to identify other unknown quantities such as spacecraft dynamics parameters, gyro biases, dynamic disturbances, or environment variations.

Lam, Quang↗

Initial Approach to Collect Small Unmanned Aircraft System Off-Nominal Operational Situations Data

NASA is developing the Unmanned Aircraft System Traffic Management research platform to safely integrate small unmanned aircraft operations in large-scale at low-altitudes. As a part of this effort, small unmanned aircraft system off-nominal operational situations data collection process has been developed to take lessons learned and to reinforce operational compliance. In this paper, descriptions of variables used for digital data collection and an online report form for collection of observational data from the operators (contextual data) are provided. They are used to collect off-nominal data from the Unmanned Aircraft System Traffic Management National Campaign in 2017. The digital data show that 2 out of 118 campaign operations (1.7%) encountered loss of navigation. Since the campaign aircraft used Global Positioning System for navigation, it is likely that unobstructed view of the sky at the campaign locations contributed to this small number. Also, 4 out of 47 operations (8.5%) encountered loss of communications. A relatively short distance between ground control system and aircraft, ranging from 2300 feet to 4200 feet, likely contributed to this small number. There was no data to identify the loss of communications condition, aircraft received signal strength, for the remaining 71 operations suggesting that some operators may not be monitoring unmanned aircraft communications system performance or monitoring it with different parameters. For the contextual data, due to the low number of total reports during the campaign, no significant trends emerged. This is an initial attempt to collect contextual data from small unmanned aircraft operators about off-nominal situations, and changes will be made to the future data collection to improve the amount and quality of the information.

unmanned aviation systems traffic management (UTM)↗

Initial Approach to Collect Small Unmanned Aircraft System Off-Nominal Operational Situations Data

NASA is developing the Unmanned Aircraft System Traffic Management research platform to safely integrate small unmanned aircraft operations in large-scale at low-altitudes. As a part of this effort, small unmanned aircraft system off-nominal operational situations data collection process has been developed to take lessons learned and to reinforce operational compliance. In this paper, descriptions of variables used for digital data collection and an online report form for collection of observational data from the operators (contextual data) are provided. They are used to collect off-nominal data from the Unmanned Aircraft System Traffic Management National Campaign in 2017. The digital data show that 2 out of 118 campaign operations (1.7%) encountered loss of navigation. Since the campaign aircraft used Global Positioning System for navigation, it is likely that unobstructed view of the sky at the campaign locations contributed to this small number. Also, 4 out of 47 operations (8.5%) encountered loss of communications. A relatively short distance between ground control system and aircraft, ranging from 2300 feet to 4200 feet, likely contributed to this small number. There was no data to identify the loss of communications condition, aircraft received signal strength, for the remaining 71 operations suggesting that some operators may not be monitoring unmanned aircraft communications system performance or monitoring it with different parameters. For the contextual data, due to the low number of total reports during the campaign, no significant trends emerged. This is an initial attempt to collect contextual data from small unmanned aircraft operators about off-nominal situations, and changes will be made to the future data collection to improve the amount and quality of the information.

Jung, Jaewoo↗

Periodic shock with added clock.

Periodic shock in rats maintained on variable-interval schedule of reinforcement interspersed with sequence of three different stimulus conditions

CONDITIONING (LEARNING)↗

Using rewards and penalties to obtain desired subject performance

Operant conditioning procedures, specifically the use of negative reinforcement, in achieving stable learning behavior is described. The critical tracking test (CTT) a method of detecting human operator impairment was tested. A pass level is set for each subject, based on that subject's asymptotic skill level while sober. It is critical that complete training take place before the individualized pass level is set in order that the impairment can be detected. The results provide a more general basis for the application of reward/penalty structures in manual control research.

Cook, M.↗

Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments

In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.

59 BASIC BIOLOGICAL SCIENCES↗

Pavlovian, Skinner, and Other Behaviourists' Contributions to AI

A version of the definition of intelligent behaviour will be supplied in the context of real and artificial systems. Short presentation of principles of learning, starting with Pavlovian s classical conditioning through reinforced response and operant conditioning of Thorndike and Skinner and finishing with cognitive learning of Tolman and Bandura will be given. The most important figures within behaviourism, especially those with contribution to AI, will be described. Some tools of artificial intelligence that act according to those principles will be presented. An attempt will be made to show when some simple rules for behaviour modifications can lead to a complex intelligent behaviour.

Kosinski, Withold↗

Rapid motor learning in the translational vestibulo-ocular reflex

Motor learning was induced in the translational vestibulo-ocular reflex (TVOR) when monkeys were repeatedly subjected to a brief (0.5 sec) head translation while they tried to maintain binocular fixation on a visual target for juice rewards. If the target was world-fixed, the initial eye speed of the TVOR gradually increased; if the target was head-fixed, the initial eye speed of the TVOR gradually decreased. The rate of learning acquisition was very rapid, with a time constant of approximately 100 trials, which was equivalent to <1 min of accumulated stimulation. These learned changes were consolidated over >or=1 d without any reinforcement, indicating induction of long-term synaptic plasticity. Although the learning generalized to targets with different viewing distances and to head translations with different accelerations, it was highly specific for the particular combination of head motion and evoked eye movement associated with the training. For example, it was specific to the modality of the stimulus (translation vs rotation) and the direction of the evoked eye movement in the training. Furthermore, when one eye was aligned with the heading direction so that it remained motionless during training, learning was not expressed in this eye, but only in the other nonaligned eye. These specificities show that the learning sites are neither in the sensory nor the motor limb of the reflex but in the sensory-motor transformation stage of the reflex. The dependence of the learning on both head motion and evoked eye movement suggests that Hebbian learning may be one of the underlying cellular mechanisms.

Non-NASA Center↗

Program Helps Simulate Neural Networks

Neural Network Environment on Transputer System (NNETS) computer program provides users high degree of flexibility in creating and manipulating wide variety of neural-network topologies at processing speeds not found in conventional computing environments. Supports back-propagation and back-propagation-related algorithms. Back-propagation algorithm used is implementation of Rumelhart's generalized delta rule. NNETS developed on INMOS Transputer(R). Predefines back-propagation network, Jordan network, and reinforcement network to assist users in learning and defining own networks. Also enables users to configure other neural-network paradigms from NNETS basic architecture. Small portion of software written in OCCAM(R) language.

Villarreal, James↗

Assessing the Effects of Momentary Priming on Memory Retention During an Interference Task

A memory aid, that used brief (33ms) presentations of previously learned information (target words), was assessed on its ability to reinforce memory for target words while the subject was performing an interference task. The interference task required subjects to learn new words and thus interfered with their memory of the target words. The brief presentation (momentary memory priming) was hypothesized to refresh the subjects memory of the target words. 143 subjects, in a within subject design, were given a 33ms presentation of the target memory words during the interference task in a treatment condition and a blank 33ms presentation in the control condition. The primary dependent measure, memory loss over the interference trial, was not significantly different between the two conditions. The memory prime did not appear to hinder the subjects performance on the interference task. This paper describes the experiment and the results along with suggestions for future research.

Schutte, Paul C.↗

Modeling the behavioral substrates of associate learning and memory - Adaptive neural models

Three adaptive single-neuron models based on neural analogies of behavior modification episodes are proposed, which attempt to bridge the gap between psychology and neurophysiology. The proposed models capture the predictive nature of Pavlovian conditioning, which is essential to the theory of adaptive/learning systems. The models learn to anticipate the occurrence of a conditioned response before the presence of a reinforcing stimulus when training is complete. Furthermore, each model can find the most nonredundant and earliest predictor of reinforcement. The behavior of the models accounts for several aspects of basic animal learning phenomena in Pavlovian conditioning beyond previous related models. Computer simulations show how well the models fit empirical data from various animal learning paradigms.

Lee, Chuen-Chien↗

Physics-Informed Neural Network (PINN) Prediction of Mixed Mass-Heat-Crystallization Limited Methane Hydrate Formation and Dissociation in Micro-Confinement

The creation and use of Physics-Informed Neural Networks (PINNs) for simulating the dynamics of methane hydrate formation and dissociation will be presented. The PINN framework's main benefit is its capacity to impose physical consistency with only a partial comprehension of the governing equations. This makes the algorithm especially useful for systems with little experimental evidence or a lack of theoretical knowledge. A strong basis for forecasting methane hydrate behavior over the verified operating ranges of 30.0-80.9 bar pressure and 1.0-4.0 K sub-cooling conditions is provided by the combination of conductive heat transfer equations and mixed mass-transfer–crystallization kinetics. PINNs were more accurate at predicting the mixed mass-heat-crystallization limited kinetics than conventional Artificial Neural Networks (ANNs), demonstrating remarkable predictive accuracy for methane hydrate production over the ANN model. The efficiency of incorporating physical limitations from first principles into machine learning frameworks for methane hydrate crystallizations is reinforced by these findings. For hydrate-related applications in energy generation, carbon sequestration, and climate modelling, our study establishes PINNs as a computational tool that is both scalable and efficient. The proven capacity to close the gap between conventional physics-based simulations and solely data-driven models creates new opportunities for expedited hydrate research and practical applications.

Hartman, Ryan L [NYU Tandon School of Engineering]↗