Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Multi-Agent Methods for the Configuration of Random Nanocomputers

As computational devices continue to shrink, the cost of manufacturing such devices is expected to grow exponentially. One alternative to the costly, detailed design and assembly of conventional computers is to place the nano-electronic components randomly on a chip. The price for such a trivial assembly process is that the resulting chip would not be programmable by conventional means. In this work, we show that such random nanocomputers can be adaptively programmed using multi-agent methods. This is accomplished through the optimization of an associated high dimensional error function. By representing each of the independent variables as a reinforcement learning agent, we are able to achieve convergence must faster than with other methods, including simulated annealing. Standard combinational logic circuits such as adders and multipliers are implemented in a straightforward manner. In addition, we show that the intrinsic flexibility of these adaptive methods allows the random computers to be reconfigured easily, making them reusable. Recovery from faults is also demonstrated.

Lawson, John W.↗

Adaptive, Distributed Control of Constrained Multi-Agent Systems

Product Distribution (PO) theory was recently developed as a broad framework for analyzing and optimizing distributed systems. Here we demonstrate its use for adaptive distributed control of Multi-Agent Systems (MASS), i.e., for distributed stochastic optimization using MAS s. First we review one motivation of PD theory, as the information-theoretic extension of conventional full-rationality game theory to the case of bounded rational agents. In this extension the equilibrium of the game is the optimizer of a Lagrangian of the (Probability dist&&on on the joint state of the agents. When the game in question is a team game with constraints, that equilibrium optimizes the expected value of the team game utility, subject to those constraints. One common way to find that equilibrium is to have each agent run a Reinforcement Learning (E) algorithm. PD theory reveals this to be a particular type of search algorithm for minimizing the Lagrangian. Typically that algorithm i s quite inefficient. A more principled alternative is to use a variant of Newton's method to minimize the Lagrangian. Here we compare this alternative to RL-based search in three sets of computer experiments. These are the N Queen s problem and bin-packing problem from the optimization literature, and the Bar problem from the distributed RL literature. Our results confirm that the PD-theory-based approach outperforms the RL-based scheme in all three domains.

Bieniawski, Stefan↗

Product Distribution Theory for Control of Multi-Agent Systems

Product Distribution (PD) theory is a new framework for controlling Multi-Agent Systems (MAS's). First we review one motivation of PD theory, as the information-theoretic extension of conventional full-rationality game theory to the case of bounded rational agents. In this extension the equilibrium of the game is the optimizer of a Lagrangian of the (probability distribution of) the joint stare of the agents. Accordingly we can consider a team game in which the shared utility is a performance measure of the behavior of the MAS. For such a scenario the game is at equilibrium - the Lagrangian is optimized - when the joint distribution of the agents optimizes the system's expected performance. One common way to find that equilibrium is to have each agent run a reinforcement learning algorithm. Here we investigate the alternative of exploiting PD theory to run gradient descent on the Lagrangian. We present computer experiments validating some of the predictions of PD theory for how best to do that gradient descent. We also demonstrate how PD theory can improve performance even when we are not allowed to rerun the MAS from different initial conditions, a requirement implicit in some previous work.

Lee, Chia Fan↗

Coordination in Large Collectives

Finding the subset of a set of imperfect devices (e.g., nano or micro devices) that results in the best aggregate device is a challenging problem. It is an abstraction of what will likely be a major difficulty in designing and controlling systems of nano or micro-scale components, particularly when a large fraction of those components may be unreliable. Rather than approaching this as a passive search problem, we transform the problem into one of coordination in a complex system by imbuing each device with simple decision making ability. In doing so we face the challenge of determining what each component should attempt to do so that the collective behavior solves the overall problem. Furthermore, in this padicular instance, we face problems of scaling (number of components in the thousands to tens of thousands), observability (components have limited sensing capabilities), and reliability (the components are faulty). We present an approach based on deriving component goals that are aligned with the overall system goal (e.g., forming best aggregate device), and can be computed using information readily (e.g., locally) available to the components. Then, each component in such a collective uses a simple reinforcement learning algorithm to selfishly pursue its own goals. Because those goals are derived in a principled manner, there is no need to use external mechanisms to force collaboration or coordination among the components to ensure that the system reaches a globally desirable solution. The results show that not only this approach provides improvements of over an order of magnitude over both traditional search methods and traditional multi-agent methods, but that the gains increase with the size of the system. This latter result makes this method ideal for domains where the number of components is currently in the thousands and will reach millions in the near future.

Tumer, Kagan↗

The Design of Collectives of Agents to Control Non-Markovian Systems

The Collective Intelligence (COIN) framework concerns the design of collectives of reinforcement-learning agents such that their interaction causes a provided "world" utility function concerning the entire collective to be maximized. Previously, we applied that framework to scenarios involving Markovian dynamics where no re-evolution of the system from counter-factual initial conditions (an often expensive calculation) is permitted. This approach sets the individual utility function of each agent to be both aligned with the world utility, and at the same time, easy for the associated agents to optimize. Here we extend that approach to systems involving non-Markovian dynamics. In computer simulations, we compare our techniques with each other and with conventional "team games". We show whereas in team games performance often degrades badly with time, it steadily improves when our techniques are used. We also investigate situations where the system's dimensionality is effectively reduced. We show that this leads to difficulties in the agents ability to learn. The implication is that learning is a property only of high-enough dimensional systems.

Lawson, John W.↗

A Decision-Theoretic Approach to Autonomous Planetary Rover Control

The report discusses the: Decentralized Control of Markov Decision Processes. Study the complexity of decentralized control of Markov decision processes, and develop algorithms for finding optimal control policies. Scheduling Contract Algorithms. Develop an optimal method for scheduling runs of a contract anytime algorithm (one that takes the deadline as input) in situations where the deadline is unknown, multiple problem instances must be solved, and a multi-processor machine is available. Planetary Rover Control as a Markov Decision Process.Use the Markov decision process framework to formalize and solve problems in planetary rover control. Adaptive Peer Selection. Use reinforcement learning to maximize the expected down-load speed for a client in a peer-to-peer file sharing system.

Zilberstein, Shlomo↗

Distributed Adaptive Control: Beyond Single-Instant, Discrete Variables

In extensive form noncooperative game theory, at each instant t, each agent i sets its state x, independently of the other agents, by sampling an associated distribution, q(sub i)(x(sub i)). The coupling between the agents arises in the joint evolution of those distributions. Distributed control problems can be cast the same way. In those problems the system designer sets aspects of the joint evolution of the distributions to try to optimize the goal for the overall system. Now information theory tells us what the separate q(sub i) of the agents are most likely to be if the system were to have a particular expected value of the objective function G(x(sub 1),x(sub 2), ...). So one can view the job of the system designer as speeding an iterative process. Each step of that process starts with a specified value of E(G), and the convergence of the q(sub i) to the most likely set of distributions consistent with that value. After this the target value for E(sub q)(G) is lowered, and then the process repeats. Previous work has elaborated many schemes for implementing this process when the underlying variables x(sub i) all have a finite number of possible values and G does not extend to multiple instants in time. That work also is based on a fixed mapping from agents to control devices, so that the the statistical independence of the agents' moves means independence of the device states. This paper also extends that work to relax all of these restrictions. This extends the applicability of that work to include continuous spaces and Reinforcement Learning. This paper also elaborates how some of that earlier work can be viewed as a first-principles justification of evolution-based search algorithms.

Wolpert, David H.↗

Designing Agent Utilities for Coordinated, Scalable and Robust Multi-Agent Systems

Coordinating the behavior of a large number of agents to achieve a system level goal poses unique design challenges. In particular, problems of scaling (number of agents in the thousands to tens of thousands), observability (agents have limited sensing capabilities), and robustness (the agents are unreliable) make it impossible to simply apply methods developed for small multi-agent systems composed of reliable agents. To address these problems, we present an approach based on deriving agent goals that are aligned with the overall system goal, and can be computed using information readily available to the agents. Then, each agent uses a simple reinforcement learning algorithm to pursue its own goals. Because of the way in which those goals are derived, there is no need to use difficult to scale external mechanisms to force collaboration or coordination among the agents, or to ensure that agents actively attempt to appropriate the tasks of agents that suffered failures. To present these results in a concrete setting, we focus on the problem of finding the sub-set of a set of imperfect devices that results in the best aggregate device. This is a large distributed agent coordination problem where each agent (e.g., device) needs to determine whether to be part of the aggregate device. Our results show that the approach proposed in this work provides improvements of over an order of magnitude over both traditional search methods and traditional multi-agent methods. Furthermore, the results show that even in extreme cases of agent failures (i.e., half the agents failed midway through the simulation) the system's performance degrades gracefully and still outperforms a failure-free and centralized search algorithm. The results also show that the gains increase as the size of the system (e.g., number of agents) increases. This latter result is particularly encouraging and suggests that this method is ideally suited for domains where the number of agents is currently in the thousands and will reach tens or hundreds of thousands in the near future.

Tumer, Kagan↗

Towards Intelligent Control for Next Generation Aircraft

NASA Aeronautics Subsonic Fixed Wing Project is focused on mitigating the environmental and operation impacts expected as aviation operations triple by 2025. The approach is to extend technological capabilities and explore novel civil transport configurations that reduce noise, emissions, fuel consumption and field length. Two Next Generation (NextGen) aircraft have been identified to meet the Subsonic Fixed Wing Project goals - these are the Hybrid Wing-Body (HWB) and Cruise Efficient Short Take-Off and Landing (CESTOL) aircraft. The technologies and concepts developed for these aircraft complicate the vehicle s design and operation. In this paper, flight control challenges for NextGen aircraft are described. The objective of this paper is to examine the potential of state-of-the-art control architectures and algorithms to meet the challenges and needed performance metrics for NextGen flight control. A broad range of conventional and intelligent control approaches are considered, including dynamic inversion control, integrated flight-propulsion control, control allocation, adaptive dynamic inversion control, data-based predictive control and reinforcement learning control.

Acosta, Diana Michelle↗

Differential Adaptive Stress Testing of Airborne Collision Avoidance Systems

The next-generation Airborne Collision Avoidance System (ACAS X) is currently being developed and tested to replace the Traffic Alert and Collision Avoidance System (TCAS) as the next international standard for collision avoidance. To validate the safety of the system, stress testing in simulation is one of several approaches for analyzing near mid-air collisions (NMACs). Understanding how NMACs can occur is important for characterizing risk and informingdevelopment of the system. Recently, adaptive stress testing (AST) has been proposed as a way to find the most likely path to a failure event. The simulation-based approach accelerates search by formulating stress testing as a sequential decision process then optimizing it using reinforcement learning. The approach has been successfully applied to stress test a prototype of ACAS Xin various simulated aircraft encounters. In some applications, we are not as interestedin the system's absolute performance as its performance relative to another system. Such situations arise, for example, during regression testing or when deciding whether a new system should replace an existing system. In our collision avoidance application, we are interested in finding cases where ACAS X fails but TCAS succeeds in resolving a conflict. Existing approaches do not provide an efficient means to perform this type of analysis. This paper extends the AST approach to differential analysis by searching two simulators simultaneously and maximizing the difference between their outcomes. We call this approach differential adaptive stress testing (DAST). We apply DAST to compare a prototype of ACAS X against TCAS and show examples of encounters found by the algorithm.

Lee, Ritchie↗

Smart Congestion Control for Delay- and Disruption Tolerant Networks

In this paper, we propose a novel congestion control framework for delay- and disruption tolerant networks (DTNs). The proposed framework, called Smart-DTN-CC, adjusts its operation automatically as a function of the dynamics of the underlying network. It employs reinforcement learning, a machine learning technique known to be well suited to problems in which the environment, in this case the network, plays a crucial role; yet, no prior knowledge about the target environment can be assumed, i.e., the only way to acquire information about the environment is to interact with it through continuous online learning. Smart-DTN-CC nodes get input from the environment (e.g., its buffer occupancy, set of neighbors, etc), and, based on that information, choose an action to take from a set of possible actions. Depending on an action’s effectiveness in controlling congestion, it will be given a reward. Smart-DTN-CC’s goal is to maximize the overall reward which translates to minimizing congestion. To our knowledge, Smart-DTN-CC is the first DTN congestion control framework that has the ability to automatically and continuously adapt to the dynamics of the target environment which allows Smart-DTNCC to deliver adequate performance in a variety of DTN applications and scenarios. As demonstrated by our experimental evaluation, Smart-DTN-CC is able to consistently outperform existing DTN congestion control mechanisms under a wide range of network conditions and characteristics.

Hirata, Celso M.↗

Piezoelectric impedance-based high-accuracy damage identification using sparsity conscious multi-objective optimization inverse analysis

Two elements are essential in structural health monitoring utilizing dynamic responses: response measurement with high-frequency contents, i.e., small characteristic wavelengths, that can adequately reflect damage features, and effective inverse identification analysis that is however oftentimes under-determined. The advancement of smart structure integration has led to active interrogation through frequency-sweeping piezoelectric impedance measurement at high frequency range. In this research we develop a multi-objective optimization formulation for the identification of damage location and severity utilizing piezoelectric impedance. While one optimization objective is to match the response measurement with finite element model prediction in the damage parametric space, the other is the number of locations of damage, i.e., the sparsity of damage index as the solution vector, since damage usually occurs within a small number of locations. This multi-objective formulation fits well the under-determined nature of damage identification, as it naturally provides multiple solutions as basis for further elucidation. The challenge remaining is how to find a small solution set that can include the actual damage scenario. Here we develop a novel inverse identification framework utilizing the intelligent swarm optimizer which possesses flexibility for enhancement. We first embed a sparsity enforcement process into the population generation of the optimizer, which yields a solution repository intrinsically possessing sparsity. We then apply reinforcement learning so the agents can adaptively opt for local strategies with the aim of enriching the searching patterns to diversify the solutions. Through the incorporation of a Q-table, searching toward more promising directions will be rewarded. Our case analyses employing experimental data indicate that this sparsity-conscious multi-objective particle swarm optimization technique can lead to a small solution set which generally encompasses the true damage scenario. This effectively solves the structural damage identification problem with piezoelectric impedance measurement.

Yang Zhang↗

Transient Stability Enhancement via a Scalable RL Method with VSG Parameter Tuning

This paper presents a reinforcement learning (RL)-driven strategy to improve the transient stability of power systems via tuning parameters of multiple virtual synchronous generators (VSGs). We proposed a scalable method to support RL training convergence probability and speed, even when a large number of contingencies are considered. The proposed scalable RL framework first decomposes the large number of contingencies into multiple groups and then conducts parallel training for each group, decreasing the state space and complexity of each training. Additionally, we propose a contingency grouping algorithm to streamline the RL action space and facilitate the training. The proposed method is validated across various standard test systems.

Huang, Xiaoge↗

Digital Twin Framework for PIP-II Linac: AI-Driven Multi-Scale Modeling from Ion Source to 800 MeV

The PIP-II linac will enable >1.2 MW beam power for DUNE, requiring unprecedented operational reliability across its warm front-end (RFQ, MEBT) and five distinct SRF sections operating at 162.5/325/650 MHz. We present a comprehensive digital twin framework uniquely combining a fully differentiable fast beam transport code with neural network surrogates trained on high-fidelity PIC simulations, capturing space charge and nonlinear dynamics beyond traditional envelope codes while achieving 10⁴× speedup at <1% accuracy. End-to-end differentiability enables gradient-based optimization across 500+ parameters simultaneously—previously impossible with conventional tools—while the model incorporates static/dynamic errors and serves as a virtual commissioning platform for diverse hardware integration. The framework facilitates reinforcement learning for pulsed/CW mode transitions, predictive maintenance through anomaly detection, and autonomous tuning algorithm development with real-time execution capability. Validation against physics simulations shows excellent agreement for the front-end, with initial results demonstrating potential for 30% commissioning time reduction and proactive fault mitigation, providing a scalable blueprint for operating next-generation high-intensity accelerators.

Pathak, Abhishek [Fermilab] (ORCID:000000021704208↗

Intern Poster Session 08/13: Autonomous Nuclear Robotics: Applications in nuclear waste inspection and hot cell experiments

The nuclear industry is experiencing renewed interest in autonomous robotics, yet most deployed systems remain teleoperated with limited autonomy. This work presents two contributions toward fully autonomous nuclear robotic systems: autonomous waste inspection at the Hanford Site and an autonomous hot cell laboratory framework. Inspections of Hanford's underground waste storage tanks are performed manually at significant cost and personnel exposure. We developed a reinforcement-learning (RL) training pipeline for a custom-built inspection arm. In parallel, we are designing an autonomous laboratory framework for post-irradiation examination in hot cells at the Specimen Preparation Laboratory (SPL) that integrates computer vision, task and motion planning, hardware execution, and operator-in-the-loop control. These systems demonstrate a path toward safer, more efficient nuclear operations by reducing human exposure while maintaining rigorous human oversight at critical decision points.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

CTRL-STEER: Closed-Loop Neuron Activation Control in Vision-Language-Action Models

Vision-Language-Action (VLA) models enable test-time behavioral steering via neuron-level interventions, but existing methods use fixed strengths and operate in open loop. This static modulation fails under evolving task dynamics, leading to overcorrection, oscillations, and reduced task success—especially for temporal attributes like speed. We propose CTRL-STEER, a control-theoretic framework that casts activation steering as closed-loop feedback with adaptive, time-varying interventions. Instead of assuming neurons encode temporal concepts, we steer along motion-aligned residual directions and regulate intervention magnitude via feedback. We instantiate this with both PID and reinforcement learning controllers that jointly optimize concept adherence and task success. Experiments on fine-tuned OpenVLA policies across four LIBERO suites show improved stability and a better steering–success trade-off over fixed-coefficient baselines, without retraining the base model.

Babu, Abhijith [Florida International University, ↗

Resilient information and inference networks under mixed-trust sensing

With ubiquitous digitization, sensing, and computational intelligence deployed in increasingly more and broader domains, including critical infrastructure, potentially misleading and destabilizing effects of multimodal anomalies and adversarial behavior are growing in importance. Here, we develop randomized and reinforcement learning-based strategies for strategically recruiting and utilizing deployed (and, thus, vulnerable and potentially faulty and/or compromised) nodes from information and inference networks, while defending against adversaries that attempt to misguide assessments of inferred variables. Recognizing that, besides communication and other costs, sampling from any observable node can either provide true data or dangerously expose our inference to misinformation (without being easily distinguishable what actually happens), the proposed strategies proceed by progressively recruiting nodes and cautiously scaling their information contribution based on assumed, or, in our reinforcement learning approach, intelligently weighed trustworthiness, with the learning approach also considering network-wide, threat-inclusive risk/value tradeoffs. While avoiding the hardware, communication, analytical and computational burden of explicit redundancy, the proposed defensive schemes enable on-the-fly assessments of underlying processes, and system-wide situational awareness with demonstrable resilience against adversarial activities.

97 - MATHEMATICS AND COMPUTING↗