Search NASASearch

SEARCH · Search NASA

Results for “proximal policy optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

29 records · Page 2

Scheduling the NASA Deep Space Network with Deep Reinforcement Learning

With three complexes spread evenly across the Earth, NASA’s Deep Space Network (DSN) is the primary means of communications as well as a significant scientific instrument for dozens of active missions around the world. A rapidly rising number of spacecraft and increasingly complex scientific instruments with higher bandwidth requirements have resulted in demand that exceeds the network’s capacity across its 12 antennae. The existing DSN scheduling process operates on a rolling weekly basis and is time-consuming; for a given week, generation of the final baseline schedule of spacecraft tracking passes takes roughly 5 months from the initial requirements submission deadline, with several weeks of peer-to-peer negotiations in between. This paper proposes a deep reinforcement learning (RL) approach to generate candidate DSN schedules from mission requests and spacecraft ephemeris data with demonstrated capability to address real-world operational constraints. A deep RL agent is developed that takes mission requests for a given week as input, and interacts with a DSN scheduling environment to allocate tracks such that its reward signal is maximized. A comparison is made between an agent trained using Proximal Policy Optimization and its random, untrained counterpart. The results represent a proof-of-concept that, given a well-shaped reward signal, a deep RL agent can learn the complex heuristics used by experts to schedule the DSN. A trained agent can potentially be used to generate candidate schedules to bootstrap the scheduling process and thus reduce the turnaround cycle for DSN scheduling.

Wilson, Brian

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L.

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L

Exploring the Low-Thrust Transfer Design Space in an Ephemeris Model via Multi-Objective Reinforcement Learning

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to train multiple policies to uncover solutions within a multi-objective solution space. MRPPO is used in this paper to train policies to construct low-thrust transfers for a SmallSat from the vicinity of 𝐿2 to an 𝐿5 short period orbit in the Sun-Earth-Moon system. First, the policies are trained in this scenario in the circular restricted three-body problem. This information is used to initialize the policies before training in a higher-fidelity ephemeris model; a process known as transfer learning. The recovered segments of the solution space will be compared to fundamental dynamical structures to both examine the results of MRPPO in this complex design scenario and explore the effectiveness of transfer learning.

Mashiku, Alinda K.

Reinforcement Learning Approach to Flight Control Allocation with Distributed Electric Propulsion

The flight control system of the SUSAN Electrofan concept aircraft achieves attitude control using both conventional flight control surfaces and differential thrust through distributed electric propulsion (DEP) from sixteen wing-mounted electric engines. The introduction of eight pairs of wing fans for attitude control creates a highly actuated system. Such a system requires more sophisticated control to operate, especially in the presence of wingfan failures where the loss of a single wingfan can result in a thrust imbalance. This paper investigates the use of deep reinforcement learning (RL) using proximal policy optimization (PPO) to achieve attitude control through a combination of DEP and control surface deflections. First, the paper examines the aircraft undergoing a coordinated turn. Then, it examines the aircraft experiencing a wingfan failure during cruise conditions. It is shown that deep reinforcement learning can be a potential avenue for nonlinear flight control design.

Distributed Electric Propulsion

Nuclear microreactor transient and load-following control with deep reinforcement learning

The economic feasibility of nuclear microreactors will depend on minimizing operating costs through advancements in autonomous control, especially when these microreactors are operating alongside other types of energy systems (e.g., renewable energy). This study explores the application of deep reinforcement learning (RL) for real-time drum control in microreactors, exploring performance in regard to load-following scenarios. By leveraging a point kinetics model with thermal and xenon feedback, we first establish a baseline using a single-output RL agent, then compare it against a traditional proportional–integral–derivative (PID) controller. This study demonstrates that RL controllers, including both single- and multi-agent RL (MARL) frameworks, can achieve similar or even superior load-following performance as traditional PID control across a range of load-following scenarios. In short transients, the RL agent was able to reduce the tracking error rate in comparison to PID by one half to one third. Over extended 300-minute load-following scenarios in which xenon feedback becomes a dominant factor, PID maintained better accuracy, but RL still remained within a 1% error margin despite being trained only on short-duration scenarios. This highlights RL’s strong ability to generalize and extrapolate to longer, more complex transients, affording substantial reductions in training costs and reduced overfitting. Furthermore, when control was extended to multiple drums, MARL enabled independent drum control as well as maintained reactor symmetry constraints without sacrificing performance---an objective that standard single-agent RL could not learn. We also found that, as increasing levels of Gaussian noise were added to the power measurements, the RL controllers were able to maintain lower error rates than PID, and to do so with at least 10% and upwards of 150% less control effort. These findings illustrate RL's potential for autonomous nuclear reactor control, laying the groundwork for future integration into high-fidelity simulations and experimental validation efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Fe(3): An Evaluation Tool for Low-Altitude Air Traffic Operations

The concepts of unmanned aircraft system traffic management (UTM) and urban air mobility (UAM) are introducing high-density operations in low altitude airspace in closer proximity to populated areas than conventional high-altitude air traffic. The Flexible engine for Fast-time Evaluation of Flight Environments (Fe (sup 3)) provides the capability of statistically analyzing the high-density, high-fidelity, and low-altitude traffic system under numerous scenarios, such that stake holders can study impacts of factors in the low-altitude high-density traffic system and define requirements, policies, and protocols needed to support a safe yet efficient traffic system, and even assess operational risks and optimize flight schedules without conducting infeasible and cost-prohibitive flight tests that involve a large volume of aerial vehicles. This work provides an introduction to this simulation tool including its architecture and various models involved. Its performance and sample application in UAM and UTM are also presented.

Collision Avoidance

Fe3: An Evaluation Tool for Low-Altitude Air Traffic Operations

The concepts of unmanned aircraft system traffic management (UTM) and urban air mobility (UAM) are introducing high-density operations in low altitude airspace in closer proximity to populated areas than conventional high-altitude air traffic. The Flexible engine for Fast-time Evaluation of Flight Environments (Fe (sup 3)) provides the capability of statistically analyzing the high-density, high-fidelity, and low-altitude traffic system under numerous scenarios, such that stake holders can study impacts of factors in the low-altitude high-density traffic system and define requirements, policies, and protocols needed to support a safe yet efficient traffic system, and even assess operational risks and optimize flight schedules without conducting infeasible and cost-prohibitive flight tests that involve a large volume of aerial vehicles. This work provides an introduction to this simulation tool including its architecture and various models involved. Its performance and sample application in UAM and UTM are also presented.

Trajectory Modeling

Where to cool off: a geospatial framework for placement of cooling centers

Indoor cooling is essential to reduce heat stress and increase passive survivability during heatwaves. Although air conditioning (AC) is recommended for maintaining indoor thermal comfort, low- and medium-income households in the U.S. often do not own an AC and/or limit AC usage to reduce energy consumption and associated costs, thereby risking their health and safety. With the frequency and intensity of heatwaves increasing, cooling centers are considered an appropriate alternative to indoor cooling and a possible mitigation strategy to prevent adverse health impacts of heat exposure. However, these centers are limited in numbers and not always accessible. This requires (i) developing a geospatial framework using physical and social factors for optimal siting of cooling centers to meet future needs and (ii) ranking of existing and potential cooling centers (schools, libraries, religious institutions) based on their accessibility among vulnerable populations and proximity to healthcare facilities. We developed and deployed a geospatial framework based on the Multi-criteria Decision Analysis approach in five U.S. cities (Los Angeles (LA), Phoenix, Austin, Atlanta, Miami) to evaluate the effectiveness of the framework in ranking cooling centers based on accessibility and population coverage. The results revealed that (i) access to cooling centers varies across cities and 32.2–50.7% of centers are within walking distance of the most vulnerable populations, (ii) vulnerable populations exposed to Urban Heat Island (UHI) effects are more likely to experience energy burden, and (iii) about 21.2–49.4% of population with high energy burden have access to these centers. Considering that more cooling centers are needed to assist energy burdened households alleviate heat exposure impacts, the framework developed herein could be adapted to incorporate other factors (e.g. health impacts, policies) to assess site suitability of existing shelters, identify potential sites for new cooling centers, and geo-target communities where energy efficient emerging technologies could be deployed to reduce heat stress.

58 GEOSCIENCES

Preliminary Analysis of Nuclear-Powered Data Center Scenarios

This report provides a comprehensive analysis of the potential for nuclear energy to meet the growing energy demands of data centers (DCs). It evaluates the technical, economic, and socio-environmental implications of coupling Nuclear Power Plants (NPPs) with DCs, providing initial responses to several key research questions: What is the potential increased energy demand from DCs in the U.S., in the short, medium and long term? The U.S. is experiencing a rapid increase in energy demand from DCs, with projections indicating a total increase of 24-74 GWy(e) by 2028. Meeting this demand with nuclear energy would require 27–85 GWe of installed capacity. While this surge is expected to slow in the long term, the DC industry needs reliable, scalable, and clean energy sources. How much nuclear capacity can be deployed to meet DC demand and in which timeframe? Several pathways for increasing nuclear capacity were identified, including uprates, restarts of recently retired reactors, power purchase agreements with existing fleet, and new construction. Approximately 20‒28 GWe of nuclear capacity could be dedicated to DCs by the early 2030s. How much High Assay Low Enriched Uranium (HALEU) would be needed to support some nuclear deployment scenarios for DCs? Meeting the deployment targets announced by Google and Amazon for the Kairos Power Fluoride-Salt-Cooled High-Temperature Reactor or KP-FHR (~500 MWe by 2035) and the Xe-100 (~1 GWe by 2040), respectively, requires ramping up 19.75% enriched HALEU production to ~6 t/yr by 2040. What types of nuclear energy/DC coupling options exist, and what are the different benefits/challenges? Five coupling options were analyzed, ranging from grid-connected configurations to colocated, behind-the-meter setups. Key design considerations include the proximity to high- and/or medium-voltage transmission lines, the desired internal fault tolerance, and the sources of alternative/backup power during outages. Each coupling option offers unique benefits and challenges in terms of reliability, system costs, regulation, timeline, etc. A list of NPP/DC deployment scenarios was developed, considering existing or newly built NPP or DC projects. Colocated DCs with new small modular reactors or large reactors on greenfield and brownfield sites are the focus of this report. What types of reactors, especially what size, may be incentivized by DCs? Reactor sizing optimization revealed that the ideal reactor size and number of units depend on DC demand, coupling configurations defined in this report, and other economic factors. Larger reactors are preferred for high-demand DCs and grid-connected systems, while larger number of smaller reactors are better suited for DC configurations without grid backup. Which sites may be compatible with co-located nuclear-powered DCs? Siting those projects is a complicated evaluation factoring local water resources, grid connection availability and reliability, IT infrastructure, local work force, proximity to population zones, etc. For this effort greenfield and brownfield sites such as retired coal-fired plants were used to evaluate this question. This evaluation is not meant to recommend any particular site but it highlights key siting criteria and demonstrates large-scale site availability. What are the socio-economic impacts of co-located nuclear-powered DCs? Those projects generate substantial economic benefits to the local economy, particularly in urban settings. Hyperscale DCs colocated with nuclear power plants (sized around 1 GW of power) can create nearly 1,700 jobs for annual operations and more than 7,300 jobs among the supply chain and local businesses as a result of increased household spending. Rural projects also provide significant benefits, but at lower magnitudes compared to urban deployments.

22 GENERAL STUDIES OF NUCLEAR REACTORS