Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

NEXT Generation Energy Technologies for Connected and Automated On-Road Vehicles (NEXTCAR Phase I & II)

The Ohio State University’s ARPA-E NEXTCAR project was a multi-phase, multi-year research, development, and demonstration program focused on improving the energy efficiency of connected and automated vehicles (CAVs). The team developed and validated advanced vehicle motion and powertrain control algorithms that coordinate propulsion and automation systems to optimize energy use. Key technologies included Dynamic Skip Fire engine control, predictive eco-driving functions such as Eco-Approach and Departure (Eco-AND) and Eco-Adaptive Cruise Control (Eco-ACC), and powertrain-agnostic optimization frameworks for hybrid, plug-in hybrid, and battery electric vehicles. The project successfully demonstrated up to 30% energy-efficiency improvement during real-world testing at the Transportation Research Center and the American Center for Mobility. The outcomes provide a foundation for scalable, cost-effective deployment of energy-optimized CAV technologies across the automotive industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Minority University System Engineering: A Small Satellite Design Experience Held at the Jet Propulsion Laboratory During the Summer of 1996

The University of Texas at El Paso (UTEP) in conjunction with the Jet Propulsion Laboratory (JPL), North Carolina A&T and California State University of Los Angeles participated during the summer of 1996 in a prototype program known as Minority University Systems Engineering (MUSE). The program consisted of a ten week internship at JPL for students and professors of the three universities. The purpose of MUSE as set forth in the MUSE program review August 5, 1996 was for the participants to gain experience in the following areas: 1) Gain experience in a multi-disciplinary project; 2) Gain experience working in a culturally diverse atmosphere; 3) Provide field experience for students to reinforce book learning; and 4) Streamline the design process in two areas: make it more financially feasible; and make it faster.

Ordaz, Miguel Angel↗

Toward applied behavior analysis of life aloft

This article deals with systems at multiple levels, at least from cell to organization. It also deals with learning, decision making, and other behavior at multiple levels. Technological development of a human behavioral ecosystem appropriate to space environments requires an analytic and synthetic orientation, explicitly experimental in nature, dictated by scientific and pragmatic considerations, and closely approximating procedures of established effectiveness in other areas of natural science. The conceptual basis of such an approach has its roots in environmentalism which has two main features: (1) knowledge comes from experience rather than from innate ideas, divine revelation, or other obscure sources; and (2) action is governed by consequences rather than by instinct, reason, will, beliefs, attitudes or even the currently fashionable cognitions. Without an experimentally derived data base founded upon such a functional analysis of human behavior, the overgenerality of "ecological systems" approaches render them incapable of ensuring the successful establishment of enduring space habitats. Without an experimentally derived function account of individual behavioral variability, a natural science of behavior cannot exist. And without a natural science of behavior, the social sciences will necessarily remain in their current status as disciplines of less than optimal precision or utility. Such a functional analysis of human performance should provide an operational account of behavior change in a manner similar to the way in which Darwin's approach to natural selection accounted for the evolution of phylogenetic lines (i.e., in descriptive, nonteleological terms). Similarly, as Darwin's account has subsequently been shown to be consonant with information obtained at the cellular level, so too should behavior principles ultimately prove to be in accord with an account of ontogenetic adaptation at a biochemical level. It would thus seem obvious that the most productive conceptual and methodological approaches to long-term research investments focused upon human behavior in space environments will require multidisciplinary inputs from such wide-ranging fields as molecular biology, environmental physiology, behavioral biology, architecture, sociology, and political science, among others.

Review↗

UAS Conflict-Avoidance Using Multiagent RL with Abstract Strategy Type Communication

The use of unmanned aerial systems (UAS) in the national airspace is of growing interest to the research community. Safety and scalability of control algorithms are key to the successful integration of autonomous system into a human-populated airspace. In order to ensure safety while still maintaining efficient paths of travel, these algorithms must also accommodate heterogeneity of path strategies of its neighbors. We show that, using multiagent RL, we can improve the speed with which conflicts are resolved in cases with up to 80 aircraft within a section of the airspace. In addition, we show that the introduction of abstract agent strategy types to partition the state space is helpful in resolving conflicts, particularly in high congestion.

Unmanned Autonomous Systems↗

Announced Strategy Types in Multiagent RL for Conflict-Avoidance in the National Airspace

The use of unmanned aerial systems (UAS) in the national airspace is of growing interest to the research community. Safety and scalability of control algorithms are key to the successful integration of autonomous system into a human-populated airspace. In order to ensure safety while still maintaining efficient paths of travel, these algorithms must also accommodate heterogeneity of path strategies of its neighbors. We show that, using multiagent RL, we can improve the speed with which conflicts are resolved in cases with up to 80 aircraft within a section of the airspace. In addition, we show that the introduction of abstract agent strategy types to partition the state space is helpful in resolving conflicts, particularly in high congestion.

National Airspace↗

AdaStress

This is a tutorial on AdaStress, a tool for finding and analyzing the likeliest failures in a simulated system under test. The presentation outlines the adaptive stress testing framework, provides a demonstration of use, and showcases several examples of failure detection in a complex real-world system.

Reinforcement learning↗

Transient Optimization for the Betterment of Turbine Electrified Energy Management

Gas turbine engine transients are associated with degraded compressor operability, which must be addressed by the engine control system and accounted for in the engine design. Failure to do so may result in events such as compressor stall/surge and combustor blow out. Transient operability concerns constrain the engine design and can result in sacrifices of efficiency and/or thrust responsiveness. The traditional approach to transient operability management is control logic that limits the fuel flow command. A companion paper presents a strategy for optimizing the transient fuel flow control logic taking into consideration transient operability and thrust responsiveness. The study covered here extends this idea to an electrified gas turbine engine that employs a power/energy management concept known as Turbine Electrified Energy Management (TEEM). TEEM uses an electric power system interfaced with the engine (hence the term ‘electrified gas turbine engine’) to further improve transient operability and alleviate associated design constraints. There can be costs associated with implementing TEEM in terms of power and energy requirements that impact the size of the electrical power system. However, the results of this study show that through optimization of the transient limit logic, power and energy requirements needed to implement TEEM can be significantly reduced. Among the conclusions that can be drawn from the results of the illustrative application covered herein are: (1) there is a reduction in the electric machine power requirement to manage operability during accelerations by 200 to 400 hp, and (2) power transfer from the low pressure spool (LPS) to the high pressure spool (HPS) is the most effective option for improving operability during decelerations, followed by the options of only injecting power on the HPS or only extracting power from the LPS.

Turbine Electrified Energy Management↗

Transient Optimization for the Betterment of Turbine Electrified Energy Management

Gas turbine engine transients are associated with degraded compressor operability, which must be addressed by the engine control system and accounted for in the engine design. Failure to do so may result in events such as compressor stall/surge and combustor blow out. Transient operability concerns constrain the engine design and can result in sacrifices of efficiency and/or thrust responsiveness. The traditional approach to transient operability management is control logic that limits the fuel flow command. A companion paper presents a strategy for optimizing the transient fuel flow control logic taking into consideration transient operability and thrust responsiveness. The study covered here extends this idea to an electrified gas turbine engine that employs a power/energy management concept known as Turbine Electrified Energy Management (TEEM). TEEM uses an electric power system interfaced with the engine (hence the term ‘electrified gas turbine engine’) to further improve transient operability and alleviate associated design constraints. There can be costs associated with implementing TEEM in terms of power and energy requirements that impact the size of the electrical power system. However, the results of this study show that through optimization of the transient limit logic, power and energy requirements needed to implement TEEM can be significantly reduced. Among the conclusions that can be drawn from the results of the illustrative application covered herein are: (1) there is a reduction in the electric machine power requirement to manage operability during accelerations by 200 to 400 hp, and (2) power transfer from the low pressure spool (LPS) to the high pressure spool (HPS) is the most effective option for improving operability during decelerations, followed by the options of only injecting power on the HPS or only extracting power from the LPS.

transient↗

Transient Optimization for the Betterment of Turbine Electrified Energy Management

Gas turbine engine transients are associated with degraded compressor operability, which must be addressed by the engine control system and accounted for in the engine design. Failure to do so may result in events such as compressor stall/surge and combustor blow out. Transient operability concerns constrain the engine design and can result in sacrifices of efficiency and/or thrust responsiveness. The traditional approach to transient operability management is control logic that limits the fuel flow command. A companion paper presents a strategy for optimizing the transient fuel flow control logic taking into consideration transient operability and thrust responsiveness. The study covered here extends this idea to an electrified gas turbine engine that employs a power/energy management concept known as Turbine Electrified Energy Management (TEEM). TEEM uses an electric power system interfaced with the engine (hence the term ‘electrified gas turbine engine’) to further improve transient operability and alleviate associated design constraints. There can be costs associated with implementing TEEM in terms of power and energy requirements that impact the size of the electrical power system. However, the results of this study show that through optimization of the transient limit logic, power and energy requirements needed to implement TEEM can be significantly reduced. Among the conclusions that can be drawn from the results of the illustrative application covered herein are: (1) there is a reduction in the electric machine power requirement to manage operability during accelerations by 200 to 400 hp, and (2) power transfer from the low pressure spool (LPS) to the high pressure spool (HPS) is the most effective option for improving operability during decelerations, followed by the options of only injecting power on the HPS or only extracting power from the LPS.

transient↗

Route-Recapturing State-Based Horizontal Maneuver Strategy for Automated Detect-and-Avoid

This report describes a novel approach to the development of a horizontal maneuver guidance strategy for Detect-and-Avoid systems. The maneuver guidance strategy provides a directive turn action that can be automatically executed by the vehicle’s auto-pilot system, taking into account the cost of recapturing the flight plan path. Pairwise conflict scenarios with non-accelerating intruders are simulated to validate the effectiveness of the maneuver guidance strategy. Initial results suggest the strategy is more effective for faster ownship than for slower ownship, which is unable to avoid conflict in certain scenarios against fast intruders. These findings indicate this novel approach shows great potential, but improvement to its performance is necessary and will be future work.

detect-and-avoid↗

Evaluating a Cognitive Extension for the Licklider Transmission Protocol in a Spacecraft Emulation Testbed

In space communications, particularly when involving regions beyond cislunar space, the development of advanced networking solutions is essential to address the challenges posed by limited connectivity, substantial propagation delays, and radio signal variations. This study explores a data-driven approach to the Licklider Transmission Protocol (LTP), specifically focusing on dynamically adjusting the maximum payload size of segments. Prior research has emphasized the potential benefits of dynamically adjusting this parameter, introducing the concept of Cognitive LTP. This paper presents a software implementation of Cognitive LTP (CLTP) within an open-source Delay Tolerant Networking (DTN) framework, specifically the High-rate Delay Tolerant Networking (HDTN), and experimentally evaluates its performance under realistic space conditions. Leveraging the Cognitive Ground Testbed (CGT), developed by NASA GRC for spacecraft communication emulation, this study effectively bridges the gap between theoretical advancements and practical applications. By thoroughly analyzing CLTP’s functionality within the CGT, this research offers insights into the practical implications of adaptive networking strategies, emphasizing the importance of conducting tests in relevant environments for the maturation of space communication technologies.

Delay Tolerant Networking↗

Reinforcement expectation in the honeybee ( Apis mellifera ): Can downshifts in reinforcement show conditioned inhibition?

When animals learn the association of a conditioned stimulus (CS) with an unconditioned stimulus (US), later presentation of the CS invokes a representation of the US. When the expected US fails to occur, theoretical accounts predict that conditioned inhibition can accrue to any other stimuli that are associated with this change in the US. Empirical work with mammals has confirmed the existence of conditioned inhibition. But the way it is manifested, the conditions that produce it, and determining whether it is the opposite of excitatory conditioning are important considerations. Invertebrates can make valuable contributions to this literature because of the well-established conditioning protocols and access to the central nervous system (CNS) for studying neural underpinnings of behavior. Nevertheless, although conditioned inhibition has been reported, it has yet to be thoroughly investigated in invertebrates. Here, we evaluate the role of the US in producing conditioned inhibition by using proboscis extension response conditioning of the honeybee (Apis mellifera). Specifically, using variations of a “feature-negative” experimental design, we use downshifts in US intensity relative to US intensity used during initial excitatory conditioning to show that an odorant in an odor–odor mixture can become a conditioned inhibitor. We argue that some alternative interpretations to conditioned inhibition are unlikely. However, we show variation across individuals in how strongly they show conditioned inhibition, with some individuals possibly revealing a different means of learning about changes in reinforcement. We discuss how the resolution of these differences is needed to fully understand whether and how conditioned inhibition is manifested in the honeybee, and whether it can be extended to investigate how it is encoded in the CNS. It is also important for extension to other insect models. In particular, work like this will be important as more is revealed of the complexity of the insect brain from connectome projects.

60 APPLIED LIFE SCIENCES↗