Search NASASearch

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Signal Whisperers: Enhancing Wireless Reception Using DRL-Guided Reflector Arrays

This paper presents a multi-agent reinforcement learning (MARL) approach for controlling adjustable metallic reflector arrays to enhance wireless signal reception in non-line-of-sight (NLOS) scenarios. Unlike conventional reconfigurable intelligent surfaces (RIS) that require complex channel estimation, our system employs a centralized training with decentralized execution (CTDE) paradigm where individual agents corresponding to reflector segments autonomously optimize reflector element orientation in three-dimensional space using spatial intelligence based on user location information. Through extensive ray-tracing simulations with dynamic user mobility, the proposed multi-agent beam-focusing framework demonstrates substantial performance improvements over single-agent reinforcement learning baselines, while maintaining rapid adaptation to user movement within one simulation step. Comprehensive evaluation across varying user densities and reflector configurations validates system scalability and robustness. The results demonstrate the potential of learning-based approaches for adaptive wireless propagation control.

deep reinforcement learning

Transient Optimization of a Gas Turbine Engine

Gas turbine engines are the primary power plants for modern commercial aircraft. Transients prompted by significant changes in thrust or power demand are common and unavoidable. Extreme transient scenarios such as those associated with a go-around during a landing attempt are possible and must be accounted for in the design of the engine and its controller. Engine transients tend to cause a reduction in compressor operability margin, which must be addressed by the engine control system and accounted for in the engine design to prevent events such as compressor stall/surge and combustor blow out. Transient operability concerns typically lead to compromises in the engine design that sacrifice efficiency and/or limit responsiveness. Transient operability is typically managed by logic that limits the fuel flow command. If this logic is not optimized, then the potential for valuable performance could be lost. This study presents a strategy for optimizing the transient limit logic and proposes a strategy for updating the control logic over the lifespan of the engine. The results demonstrate significant improvements in transient operability. For example, of the results at sea level static conditions demonstrated a 31% reduction in the usage of the high pressure compressor operability stack during a snap acceleration transient. Furthermore, a reinforcement learning algorithm is demonstrated to modify the transient logic as the engine degrades to minimize response time while respecting a prescribed compressor operability margin limit. A simple demonstration of the reinforcement learning algorithm resulted in a thrust response time reduction of ~11.8%.

transient

Transient Optimization of a Gas Turbine Engine

Gas turbine engines are the primary power plants for modern commercial aircraft. Transients prompted by significant changes in thrust or power demand are common and unavoidable. Extreme transient scenarios such as those associated with a go-around during a landing attempt are possible and must be accounted for in the design of the engine and its controller. Engine transients tend to cause a reduction in compressor operability margin, which must be addressed by the engine control system and accounted for in the engine design to prevent events such as compressor stall/surge and combustor blow out. Transient operability concerns typically lead to compromises in the engine design that sacrifice efficiency and/or limit responsiveness. Transient operability is typically managed by logic that limits the fuel flow command. If this logic is not optimized, then the potential for valuable performance could be lost. This study presents a strategy for optimizing the transient limit logic and proposes a strategy for updating the control logic over the lifespan of the engine. The results demonstrate significant improvements in transient operability. For example, of the results at sea level static conditions demonstrated a 31% reduction in the usage of the high pressure compressor operability stack during a snap acceleration transient. Furthermore, a reinforcement learning algorithm is demonstrated to modify the transient logic as the engine degrades to minimize response time while respecting a prescribed compressor operability margin limit. A simple demonstration of the reinforcement learning algorithm resulted in a thrust response time reduction of ~11.8%.

transient

Transient Optimization of a Gas Turbine Engine

Gas turbine engines are the primary power plants for modern commercial aircraft. Transients prompted by significant changes in thrust or power demand are common and unavoidable. Extreme transient scenarios such as those associated with a go-around during a landing attempt are possible and must be accounted for in the design of the engine and its controller. Engine transients tend to cause a reduction in compressor operability margin, which must be addressed by the engine control system and accounted for in the engine design to prevent events such as compressor stall/surge and combustor blow out. Transient operability concerns typically lead to compromises in the engine design that sacrifice efficiency and/or limit responsiveness. Transient operability is typically managed by logic that limits the fuel flow command. If this logic is not optimized, then the potential for valuable performance could be lost. This study presents a strategy for optimizing the transient limit logic and proposes a strategy for updating the control logic over the lifespan of the engine. The results demonstrate significant improvements in transient operability. For example, of the results at sea level static conditions demonstrated a 31% reduction in the usage of the high pressure compressor operability stack during a snap acceleration transient. Furthermore, a reinforcement learning algorithm is demonstrated to modify the transient logic as the engine degrades to minimize response time while respecting a prescribed compressor operability margin limit. A simple demonstration of the reinforcement learning algorithm resulted in a thrust response time reduction of ~11.8%.

transient

Refining fuzzy logic controllers with machine learning

In this paper, we describe the GARIC (Generalized Approximate Reasoning-Based Intelligent Control) architecture, which learns from its past performance and modifies the labels in the fuzzy rules to improve performance. It uses fuzzy reinforcement learning which is a hybrid method of fuzzy logic and reinforcement learning. This technology can simplify and automate the application of fuzzy logic control to a variety of systems. GARIC has been applied in simulation studies of the Space Shuttle rendezvous and docking experiments. It has the potential of being applied in other aerospace systems as well as in consumer products such as appliances, cameras, and cars.

Berenji, Hamid R.

Time-Extended Payoffs for Collectives of Autonomous Agents

A collective is a set of self-interested agents which try to maximize their own utilities, along with a a well-defined, time-extended world utility function which rates the performance of the entire system. In this paper, we use theory of collectives to design time-extended payoff utilities for agents that are both aligned with the world utility, and are "learnable", i.e., the agents can readily see how their behavior affects their utility. We show that in systems where each agent aims to optimize such payoff functions, coordination arises as a byproduct of the agents selfishly pursuing their own goals. A game theoretic analysis shows that such payoff functions have the net effect of aligning the Nash equilibrium, Pareto optimal solution and world utility optimum, thus eliminating undesirable behavior such as agents working at cross-purposes. We then apply collective-based payoff functions to the token collection in a gridworld problem where agents need to optimize the aggregate value of tokens collected across an episode of finite duration (i.e., an abstracted version of rovers on Mars collecting scientifically interesting rock samples, subject to power limitations). We show that, regardless of the initial token distribution, reinforcement learning agents using collective-based payoff functions significantly outperform both natural extensions of single agent algorithms and global reinforcement learning solutions based on "team games".

Tumer, Kagan

A Dynamic Pricing Method to Manage the Impact of EV Charging on the Grid Using RL

This work addresses the challenge of managing electrical vehicle (EV) charging loads on distribution feeders with the increase in deployment of fast charging stations. To mitigate the adverse impacts on feeder health, a novel dynamic grid-informed pricing approach is proposed. This approach leverages reinforcement learning (RL) to determine hourly charging prices based on real-time grid conditions. A synthetic environment was developed to train the reinforcement learning agent. A model of an IEEE 34-bus distribution feeder with EV charging stations has been developed in OpenDSS utilizing Caldera for realistic EV charging profiles. Test cases demonstrate that the dynamic pricing strategy achieves higher energy delivery to the EV end user at a lower cost compared to constant pricing methods, while lowering voltage deviations and congestion. This approach offers more granular price adjustments, responding dynamically to feeder conditions and potentially improving grid stability and efficiency. The communication architecture to implement this dynamic pricing method is described. This research contributes to the development of smart grid-informed charging solutions that can reduce the cost of charging to the end user while also helping the grid.

EV charging, dynamic pricing, grid-informed chargi

A dynamic pricing method to manage the impact of EV charging on the grid using RL

This work addresses the challenge of managing electrical vehicle (EV) charging loads on distribution feeders with the increase in deployment of fast charging stations. To mitigate the adverse impacts on feeder health, a novel dynamic grid-informed pricing approach is proposed. This approach leverages reinforcement learning (RL) to determine hourly charging prices based on real-time grid conditions. A synthetic environment was developed to train the reinforcement learning agent. A model of an IEEE 34-bus distribution feeder with EV charging stations has been developed in OpenDSS utilizing Caldera for realistic EV charging profiles. Test cases demonstrate that the dynamic pricing strategy achieves higher energy delivery to the EV end user at a lower cost compared to constant pricing methods, while lowering voltage deviations and congestion. This approach offers more granular price adjustments, responding dynamically to feeder conditions and potentially improving grid stability and efficiency. The communication architecture to implement this dynamic pricing method is described. This research contributes to the development of smart grid-informed charging solutions that can reduce the cost of charging to the end user while also helping the grid.

24 - POWER TRANSMISSION AND DISTRIBUTION

blastforge

BlastForge is a python module meant to aid reinforcement learning research for geometry optimization research projects. The code is built to use PyTorch as a backend and will contain several model architectures and reinforcement learning training loops as well as helper python functions to evaluate a model’s performance during and after training. BlastForge is meant to be a small, focused python project to study moderate-complexity geometry optimization problems

Hickmann, Kyle [Los Alamos National Laboratory]

Advanced Computational Techniques for Improving Resilience of Critical Energy Infrastructure under Cyber-Physical Attacks

In this chapter, we present recent advances in improving the resilience of cyber-physical systems, especially with regards to energy systems. We provide discussions around various types of cyber-physical events that can cause disruptions and new advances in optimization, control, and reinforcement learning (RL) to deal with the challenges posed by such cyber-physical events. The presented methods range from distributed robust optimization, autonomous and coordinated control, reinforcement learning based resilient control and topology reconfiguration in Inter-System resilient control.

Nazir, Mohammad Nawaf [BATTELLE (PACIFIC NW LAB)]

ICE-RASSOR: Intelligent Capabilities Enhanced Regolith Advanced Surface Systems Operations Robot

NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU)processing. RASSOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar surface, RASSOR software and sensory systems need to be robust and maximize the information extracted from a reduced sensor payload. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We apply supervised learning using real data to estimate the soil mass collected without the need for mass flow rate monitors or other explicate sensing techniques. We also create a reduced-order simulation environment to develop autonomous trenching controllers via reinforcement learning and prototype state estimation architectures. Our initial results suggest that excavated regolith mass can be inferred within 2.9% RMS error of full scale, and reinforcement learning for autonomous operations has learned viable trenching strategies and helped identify desirable sensing capabilities, arrangements, and considerations. Future work includes regolith mass estimation during dynamic operation, expanding our simulation to more complex environments, and transfer learning from simulation to hardware.

machine learning

ICE-RASSOR: Intelligent Capabilities Enhanced

NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU) processing. RAS-SOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar sur-face, RASSOR software and sensory systems need to be robust and maximize the information extracted from on-board sensing. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We apply supervised learning using real data to estimate the soil mass collected without the need for mass flow rate monitors or other explicate sensing techniques. We also create a reduced-order simulation environment to develop autonomous trenching controllers via reinforcement learning and proto-type state estimation architectures. Our initial results suggest that excavated regolith mass can be inferred within 2.9% RMS error of full scale, and reinforcement learning for autonomous operations has learned viable trenching strategies and helped identify desirable sensing capabilities, arrangements, and considerations. Future work includes regolith mass estimation during dynamic operation, expanding our simulation to more complex environments, and transfer learning from simulation to hardware.

machine learning

ICE-RASSOR: Intelligent Capabilities Enhanced Regolith Advanced Surface Systems Operations Robot

NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU) processing. RASSOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar surface, RASSOR software and sensory systems need to be robust and maximize the information extracted from on-board sensory. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We apply supervised learning using real data to estimate the soil mass collected without the need for mass flow rate monitors or other explicate sensing techniques. We also create a reduced-order simulation environment to develop autonomous trenching controllers via reinforcement learning and proto-type state estimation architectures. Our initial results suggest that excavated regolith mass can be inferred within 2.9% RMS error of full scale, and reinforcement learning for autonomous operations has learned viable trenching strategies and helped identify desirable sensing capabilities, arrangements, and considerations. Future work includes regolith mass estimation during dynamic operation, expanding our simulation to more complex environments, and transfer learning from simulation to hardware.

machine learning

Everybody Needs Somebody Sometimes: Validation of Adaptive Recovery in Robotic Space Operations

This work assesses an adaptive approach to fault recovery in autonomous robotic space operations, which uses indicators of opportunity, such as physiological state measurements and observations of past human assistant performance, to inform future selections. We validated our reinforcement learning approach using data we collected from humans executing simulated mission scenarios. We present a method of structuring human-factors experiments that permits collection of relevant indicator of opportunity and assigned assistance task performance data, as well as evaluation of our adaptive approach, without requiring large numbers of test subjects. Application of our reinforcement learning algorithm to our experimental data shows that our adaptive assistant selection approach can achieve lower cumulative regret compared to existing non-adaptive baseline approaches when using real human data. Our work has applications beyond space robotics to any application where autonomy failures may occur that require external intervention.

Human-centered robotics

Product Distributions for Distributed Optimization

With connections to bounded rational game theory, information theory and statistical mechanics, Product Distribution (PD) theory provides a new framework for performing distributed optimization. Furthermore, PD theory extends and formalizes Collective Intelligence, thus connecting distributed optimization to distributed Reinforcement Learning (FU). This paper provides an overview of PD theory and details an algorithm for performing optimization derived from it. The approach is demonstrated on two unconstrained optimization problems, one with discrete variables and one with continuous variables. To highlight the connections between PD theory and distributed FU, the results are compared with those obtained using distributed reinforcement learning inspired optimization approaches. The inter-relationship of the techniques is discussed.

Bieniawski, Stefan R.

Machine Learning an Ab-Initio Based Bond-Order Potential for Bismuthene

Bismuthene is a heavy 2D material whose strong spin–orbit coupling and recently observed single-element ferroelectricity have intensified interest in its structural, vibrational, and transport properties. Accurate modeling of these behaviors requires a short-range interatomic potential that can reproduce the underlying bonding physics at a fraction of the computational cost of first-principles methods. However, such a potential is currently unavailable. Here, in this work, we construct a Tersoff bond-order potential for β-bismuthene using a reinforcement-learning framework that integrates a continuous Monte Carlo Tree Search with a simplex-based local optimizer. The optimized parameter sets reproduce first-principles lattice constants, cohesive energy, the equation of state, elastic constants, and phonon dispersion. We validate the models by performing thermal-conductivity calculations and uniaxial fracture simulations our findings confirm the reliability of the resulting models across multiple thermomechanical regimes. Comparison of the three best solutions reveals how differences in pairwise interactions, angular terms, and bond-order behavior govern phonon features and mechanical responses. We demonstrate an interpretable and computationally efficient potential for bismuthene and demonstrate a general reinforcement-learning strategy for developing bond-order models in emerging 2D materials.

deformation

Towards AI-assisted neutrino flavor theory design

Particle physics theories, such as those which explain neutrino flavor mixing, arise from a vast landscape of model-building possibilities. A model’s construction typically relies on the intuition of theorists. It also requires considerable effort to identify appropriate symmetry groups, assign field representations, and extract predictions for comparison with experimental data. We develop Autonomous Model Builder (AMBer), a framework in which a reinforcement learning agent interacts with a streamlined physics software pipeline to search these spaces efficiently. AMBer selects symmetry groups, particle content, and group representation assignments to construct models while minimizing the number of free parameters introduced. We validate our approach in well-studied regions of theory space and extend the exploration to a previously unexamined symmetry group. While demonstrated in the context of neutrino flavor theories, this approach of reinforcement learning with physics software feedback may be extended to other theoretical model-building problems in the future.

Baretz, Jason Benjamin

Modeling Humans as Reinforcement Learners: How to Predict Human Behavior in Multi-Stage Games

This paper introduces a novel framework for modeling interacting humans in a multi-stage game environment by combining concepts from game theory and reinforcement learning. The proposed model has the following desirable characteristics: (1) Bounded rational players, (2) strategic (i.e., players account for one anothers reward functions), and (3) is computationally feasible even on moderately large real-world systems. To do this we extend level-K reasoning to policy space to, for the first time, be able to handle multiple time steps. This allows us to decompose the problem into a series of smaller ones where we can apply standard reinforcement learning algorithms. We investigate these ideas in a cyber-battle scenario over a smart power grid and discuss the relationship between the behavior predicted by our model and what one might expect of real human defenders and attackers.

Lee, Ritchie