Search NASA⌕ Search

SEARCH · Search NASA

Results for “Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Using neural networks and Dyna algorithm for integrated planning, reacting and learning in systems

The traditional AI answer to the decision making problem for a robot is planning. However, planning is usually CPU-time consuming, depending on the availability and accuracy of a world model. The Dyna system generally described in earlier work, uses trial and error to learn a world model which is simultaneously used to plan reactions resulting in optimal action sequences. It is an attempt to integrate planning, reactive, and learning systems. The architecture of Dyna is presented. The different blocks are described. There are three main components of the system. The first is the world model used by the robot for internal world representation. The input of the world model is the current state and the action taken in the current state. The output is the corresponding reward and resulting state. The second module in the system is the policy. The policy observes the current state and outputs the action to be executed by the robot. At the beginning of program execution, the policy is stochastic and through learning progressively becomes deterministic. The policy decides upon an action according to the output of an evaluation function, which is the third module of the system. The evaluation function takes the following as input: the current state of the system, the action taken in that state, the resulting state, and a reward generated by the world which is proportional to the current distance from the goal state. Originally, the work proposed was as follows: (1) to implement a simple 2-D world where a 'robot' is navigating around obstacles, to learn the path to a goal, by using lookup tables; (2) to substitute the world model and Q estimate function Q by neural networks; and (3) to apply the algorithm to a more complex world where the use of a neural network would be fully justified. In this paper, the system design and achieved results will be described. First we implement the world model with a neural network and leave Q implemented as a look up table. Next, we use a lookup table for the world model and implement the Q function with a neural net. Time limitations prevented the combination of these two approaches. The final section discusses the results and gives clues for future work.

Lima, Pedro↗

Design issues of a reinforcement-based self-learning fuzzy controller for petrochemical process control

Fuzzy logic controllers have some often-cited advantages over conventional techniques such as PID control, including easier implementation, accommodation to natural language, and the ability to cover a wider range of operating conditions. One major obstacle that hinders the broader application of fuzzy logic controllers is the lack of a systematic way to develop and modify their rules; as a result the creation and modification of fuzzy rules often depends on trial and error or pure experimentation. One of the proposed approaches to address this issue is a self-learning fuzzy logic controller (SFLC) that uses reinforcement learning techniques to learn the desirability of states and to adjust the consequent part of its fuzzy control rules accordingly. Due to the different dynamics of the controlled processes, the performance of a self-learning fuzzy controller is highly contingent on its design. The design issue has not received sufficient attention. The issues related to the design of a SFLC for application to a petrochemical process are discussed, and its performance is compared with that of a PID and a self-tuning fuzzy logic controller.

Yen, John↗

Learning class descriptions from a data base of spectral reflectance of soil samples

Consideration is given to a program developed to learn class descriptions from positive and negative training examples of spectral reflectance data of bare soils. It is a combination of 'learning by example' and the generate-and-test paradigm and is designed to provide a robust learning environment that can handle error-prone data. The program was tested by having it learn class descriptions of various categories of organic carbon content, iron oxide content, and particle size distribution in soils. These class descriptions were then used to classify an array of targets. The program found the sequence of relationships between bands that contained the most important information to distinguish the classes. Physical explanations for the class descriptions obtained are presented.

Kimes, D. S.↗

On the asymptotic improvement of supervised learning by utilizing additional unlabeled samples - Normal mixture density case

The effect of additional unlabeled samples in improving the supervised learning process is studied in this paper. Three learning processes. supervised, unsupervised, and combined supervised-unsupervised, are compared by studying the asymptotic behavior of the estimates obtained under each process. Upper and lower bounds on the asymptotic covariance matrices are derived. It is shown that under a normal mixture density assumption for the probability density function of the feature space, the combined supervised-unsupervised learning is always superior to the supervised learning in achieving better estimates. Experimental results are provided to verify the theoretical concepts.

Shahshahani, Behzad M.↗

Learning and optimization with cascaded VLSI neural network building-block chips

To demonstrate the versatility of the building-block approach, two neural network applications were implemented on cascaded analog VLSI chips. Weights were implemented using 7-b multiplying digital-to-analog converter (MDAC) synapse circuits, with 31 x 32 and 32 x 32 synapses per chip. A novel learning algorithm compatible with analog VLSI was applied to the two-input parity problem. The algorithm combines dynamically evolving architecture with limited gradient-descent backpropagation for efficient and versatile supervised learning. To implement the learning algorithm in hardware, synapse circuits were paralleled for additional quantization levels. The hardware-in-the-loop learning system allocated 2-5 hidden neurons for parity problems. Also, a 7 x 7 assignment problem was mapped onto a cascaded 64-neuron fully connected feedback network. In 100 randomly selected problems, the network found optimal or good solutions in most cases, with settling times in the range of 7-100 microseconds.

Duong, T.↗

A new learning strategy for the two-time-scale neural controller with its application to the tracking control of rigid arms

A novel fast learning rule with fast weight identification is proposed for the two-time-scale neural controller, and a two-stage learning strategy is developed for the proposed neural controller. The results of the stability analysis show that both the tracking error and the fast weight error will be uniformly bounded and converge to a bounded region which depends only on the accuracy of the slow learning if the system is sufficiently excited. The efficiency of the two-stage learning is also demonstrated by a simulation of a two-link arm.

Cheng, W.↗

Learning a trajectory using adjoint functions and teacher forcing

A new methodology for faster supervised temporal learning in nonlinear neural networks is presented which builds upon the concept of adjoint operators to allow fast computation of the gradients of an error functional with respect to all parameters of the neural architecture, and exploits the concept of teacher forcing to incorporate information on the desired output into the activation dynamics. The importance of the initial or final time conditions for the adjoint equations is discussed. A new algorithm is presented in which the adjoint equations are solved simultaneously (i.e., forward in time) with the activation dynamics of the neural network. We also indicate how teacher forcing can be modulated in time as learning proceeds. The results obtained show that the learning time is reduced by one to two orders of magnitude with respect to previously published results, while trajectory tracking is significantly improved. The proposed methodology makes hardware implementation of temporal learning attractive for real-time applications.

Toomarian, Nikzad B.↗

Sea ice classification using fast learning neural networks

A first learning neural network approach to the classification of sea ice is presented. The fast learning (FL) neural network and a multilayer perceptron (MLP) trained with backpropagation learning (BP network) were tested on simulated data sets based on the known dominant scattering characteristics of the target class. Four classes were used in the data simulation: open water, thick lossy saline ice, thin saline ice, and multiyear ice. The BP network was unable to consistently converge to less than 25 percent error while the FL method yielded an average error of approximately 1 percent on the first iteration of training. The fast learning method presented can significantly reduce the CPU time necessary to train a neural network as well as consistently yield higher classification accuracy than BP networks.

Dawson, M. S.↗

Lift-fan aircraft: Lessons learned-the pilot's perspective

This paper is written from an engineering test pilot's point of view. Its purpose is to present lift-fan 'lessons learned' from the perspective of first-hand experience accumulated during the period 1962 through 1988 while flight testing vertical/short take-off and landing (V/STOL) experimental aircraft and evaluating piloted engineering simulations of promising V/STOL concepts. Specifically, the scope of the discussions to follow is primarily based upon a critical review of the writer's personal accounts of 30 hours of XV-5A/B and 2 hours of X-14A flight testing as well as a limited simulator evaluation of the Grumman Design 755 lift-fan aircraft. Opinions of other test pilots who flew these aircraft and the aircraft simulator are also included and supplement the writer's comments. Furthermore, the lessons learned are presented from the perspective of the writer's flying experience: 10,000 hours in 100 fixed- and rotary-wing aircraft including 330 hours in 5 experimental V/STOL research aircraft. The paper is organized to present to the reader a clear picture of lift-fan lessons learned from three distinct points of view in order to facilitate application of the lesson principles to future designs. Lessons learned are first discussed with respect to case histories of specific flight and simulator investigations. These principles are then organized and restated with respect to four selected design criteria categories in Appendix I. Lastly, Appendix Il is a discussion of the design of a hypothetical supersonic short take-off vertical landing (STOVL) fighter/attack aircraft.

Gerdes, Ronald M.↗

Establishing a Distance Learning Plan for International Space Station (ISS) Interactive Video Education Events (IVEE)

Educational outreach is an integral part of the International Space Station (ISS) mandate. In a few scant years, the International Space Station has already established a tradition of successful, general outreach activities. However, as the number of outreach events increased and began to reach school classrooms, those events came under greater scrutiny by the education community. Some of the ISS electronic field trips, while informative and helpful, did not meet the generally accepted criteria for education events, especially within the context of the classroom. To make classroom outreach events more acceptable to educators, the ISS outreach program must differentiate between communication events (meant to disseminate information to the general public) and education events (designed to facilitate student learning). In contrast to communication events, education events: are directed toward a relatively homogeneous audience who are gathered together for the purpose of learning, have specific performance objectives which the students are expected to master, include a method of assessing student performance, and include a series of structured activities that will help the students to master the desired skill(s). The core of the ISS education events is an interactive videoconference between students and ISS representatives. This interactive videoconference is to be preceded by and followed by classroom activities which help the students aftain the specified learning objectives. Using the interactive videoconference as the centerpiece of the education event lends a special excitement and allows students to ask questions about what they are learning and about the International Space Station and NASA. Whenever possible, the ISS outreach education events should be congruent with national guidelines for student achievement. ISS outreach staff should recognize that there are a number of different groups that will review the events, and that each group has different criteria for acceptance. For example, school administrators are more likely to be concerned about an event meeting national standards and the cost of the event. In contrast, a teacher's acceptance of an education event may be directly related to the amount of extra work the event imposes upon that teacher. ISS education events must be marketed differently to the different groups of educators, and must never increase the workload of the average teacher.

Wallington, Clint↗

Reinforcement Learning in a Nonstationary Environment: The El Farol Problem

This paper examines the performance of simple learning rules in a complex adaptive system based on a coordination problem modeled on the El Farol problem. The key features of the El Farol problem are that it typically involves a medium number of agents and that agents' pay-off functions have a discontinuous response to increased congestion. First we consider a single adaptive agent facing a stationary environment. We demonstrate that the simple learning rules proposed by Roth and Er'ev can be extremely sensitive to small changes in the initial conditions and that events early in a simulation can affect the performance of the rule over a relatively long time horizon. In contrast, a reinforcement learning rule based on standard practice in the computer science literature converges rapidly and robustly. The situation is reversed when multiple adaptive agents interact: the RE algorithms often converge rapidly to a stable average aggregate attendance despite the slow and erratic behavior of individual learners, while the CS based learners frequently over-attend in the early and intermediate terms. The symmetric mixed strategy equilibria is unstable: all three learning rules ultimately tend towards pure strategies or stabilize in the medium term at non-equilibrium probabilities of attendance. The brittleness of the algorithms in different contexts emphasize the importance of thorough and thoughtful examination of simulation-based results.

Bell, Ann Maria↗

A Computer Learning Center for Environmental Sciences

In the fall of 1998, MacMillan Hall opened at Brown University to students. In MacMillan Hall was the new Computer Learning Center, since named the EarthLab which was outfitted with high-end workstations and peripherals primarily focused on the use of remotely sensed and other spatial data in the environmental sciences. The NASA grant we received as part of the "Centers of Excellence in Applications of Remote Sensing to Regional and Global Integrated Environmental Assessments" was the primary source of funds to outfit this learning and research center. Since opening, we have expanded the range of learning and research opportunities and integrated a cross-campus network of disciplines who have come together to learn and use spatial data of all kinds. The EarthLab also forms a core of undergraduate, graduate, and faculty research on environmental problems that draw upon the unique perspective of remotely sensed data. Over the last two years, the Earthlab has been a center for research on the environmental impact of water resource use in and regions, impact of the green revolution on forest cover in India, the design of forest preserves in Vietnam, and detailed assessments of the utility of thermal and hyperspectral data for water quality analysis. It has also been used extensively for local environmental activities, in particular studies on the impact of lead on the health of urban children in Rhode Island. Finally, the EarthLab has also served as a key educational and analysis center for activities related to the Brown University Affiliated Research Center that is devoted to transferring university research to the private sector.

Mustard, John F.↗

On-Line, Self-Learning, Predictive Tool for Determining Payload Thermal Response

This paper will present the results of a joint ManTech / Goddard R&D effort, currently under way, to develop and test a computer based, on-line, predictive simulation model for use by facility operators to predict the thermal response of a payload during thermal vacuum testing. Thermal response was identified as an area that could benefit from the algorithms developed by Dr. Jeri for complex computer simulations. Most thermal vacuum test setups are unique since no two payloads have the same thermal properties. This requires that the operators depend on their past experiences to conduct the test which requires time for them to learn how the payload responds while at the same time limiting any risk of exceeding hot or cold temperature limits. The predictive tool being developed is intended to be used with the new Thermal Vacuum Data System (TVDS) developed at Goddard for the Thermal Vacuum Test Operations group. This model can learn the thermal response of the payload by reading a few data points from the TVDS, accepting the payload's current temperature as the initial condition for prediction. The model can then be used as a predictive tool to estimate the future payload temperatures according to a predetermined shroud temperature profile. If the error of prediction is too big, the model can be asked to re-learn the new situation on-line in real-time and give a new prediction. Based on some preliminary tests, we feel this predictive model can forecast the payload temperature of the entire test cycle within 5 degrees Celsius after it has learned 3 times during the beginning of the test. The tool will allow the operator to play "what-if' experiments to decide what is his best shroud temperature set-point control strategy. This tool will save money by minimizing guess work and optimizing transitions as well as making the testing process safer and easier to conduct.

Jen, Chian-Li↗

Quality Training and Learning in Aviation: Problems of Alignment

The challenge of producing training programs that lead to quality learning outcomes is ever present in aviation, especially when economic and regulatory pressures are brought into the equation. Previous research by Telfer & Moore (1997) indicates the importance of appropriate alignment of beliefs about learning across all levels of an organization from the managerial level, through the instructor/check and training level, to the pilots and other crew. This paper argues for a central focus on approaches to learning and training that encourage understanding, problem solving and application. Recent research in the area is emphasized as are methods and techniques for enhancing deeper learning.

Moore, Phillip J.↗

Representing Learning With Graphical Models

Probabilistic graphical models are being used widely in artificial intelligence, for instance, in diagnosis and expert systems, as a unified qualitative and quantitative framework for representing and reasoning with probabilities and independencies. Their development and use spans several fields including artificial intelligence, decision theory and statistics, and provides an important bridge between these communities. This paper shows by way of example that these models can be extended to machine learning, neural networks and knowledge discovery by representing the notion of a sample on the graphical model. Not only does this allow a flexible variety of learning problems to be represented, it also provides the means for representing the goal of learning and opens the way for the automatic development of learning algorithms from specifications.

Buntine, Wray L.↗

Learning Sequences of Actions in Collectives of Autonomous Agents

In this paper we focus on the problem of designing a collective of autonomous agents that individually learn sequences of actions such that the resultant sequence of joint actions achieves a predetermined global objective. We are particularly interested in instances of this problem where centralized control is either impossible or impractical. For single agent systems in similar domains, machine learning methods (e.g., reinforcement learners) have been successfully used. However, applying such solutions directly to multi-agent systems often proves problematic, as agents may work at cross-purposes, or have difficulty in evaluating their contribution to achievement of the global objective, or both. Accordingly, the crucial design step in multiagent systems centers on determining the private objectives of each agent so that as the agents strive for those objectives, the system reaches a good global solution. In this work we consider a version of this problem involving multiple autonomous agents in a grid world. We use concepts from collective intelligence to design goals for the agents that are 'aligned' with the global goal, and are 'learnable' in that agents can readily see how their behavior affects their utility. We show that reinforcement learning agents using those goals outperform both 'natural' extensions of single agent algorithms and global reinforcement, learning solutions based on 'team games'.

Turner, Kagan↗

NASA Materials Related Lessons Learned

Lessons Learned have been the basis for our accomplishments throughout the ages. They have been passed down from father to son, mother to daughter, teacher to pupil, and older to younger worker. Lessons Learned have also been the basis for the nation's accomplishments for more than 200 years. Both government and industry have long recognized the need to systematically document and utilize the knowledge gained from past experiences in order to avoid the repetition of failures and mishaps. Through the knowledge captured and recorded in Lessons Learned from more than 80 years of flight in the Earth's atmosphere, NASA's materials researchers are constantly working to develop stronger, lighter, and more durable materials that can withstand the challenges of space. The Agency's talented materials engineers and scientists continue to build on that rich tradition by using the knowledge and wisdom gained from past experiences to create futurist materials and technologies that will be used in the next generation of advanced spacecraft and satellites that may one day enable mankind to land men on another planet or explore our nearest star. These same materials may also have application here on Earth to make commercial aircraft more economical to build and fly. With the explosion in technical accomplishments over the last decade, the ability to capture knowledge and have the capability to rapidly communicate this knowledge at lightning speed throughout an organization like NASA has become critical. Use of Lessons Learned is a principal component of an organizational culture committed to continuous improvement.

Garcia, Danny↗

Reinforcement Learning for Weakly-Coupled MDPs and an Application to Planetary Rover Control

Weakly-coupled Markov decision processes can be decomposed into subprocesses that interact only through a small set of bottleneck states. We study a hierarchical reinforcement learning algorithm designed to take advantage of this particular type of decomposability. To test our algorithm, we use a decision-making problem faced by autonomous planetary rovers. In this problem, a Mars rover must decide which activities to perform and when to traverse between science sites in order to make the best use of its limited resources. In our experiments, the hierarchical algorithm performs better than Q-learning in the early stages of learning, but unlike Q-learning it converges to a suboptimal policy. This suggests that it may be advantageous to use the hierarchical algorithm when training time is limited.

Bernstein, Daniel S.↗