Search NASA⌕ Search

Engineering topics

John R Martin

Publications and source records attributed to John R Martin.

Reinforcement Learning for Spacecraft Navigation & Environment Characterization in the Planar-Restricted Two-Body Problem

During mission planning and execution, spacecraft operators must balance data collection and downlink, systems constraints, human factors, and navigation. As missions become increasingly complex and ambitious, these factors become more intricately entwined and conflicted. For example, a spacecraft’s position must be known accurately in order to point to and image a target. Large position errors may cause missed observations or require additional scanning that increases operations complexity and data volume. Some observations require imaging from specific relative geometries which adds orbit control and timing considerations. Adjusting the orbit may allow for optimal observability of environmental parameters and/or enable more efficient sensor coverage, but maneuver execution error adds uncertainty to the current state which impacts both characterization and coverage objectives.

Navigation↗

Reinforcement Learning for Spacecraft Navigation & Environment Characterization in the Planar-Restricted Two-Body Problem

As science, exploration, and commercial space missions become increasingly complex, so does the need for efficient, autonomous, and integrated spacecraft navigation and operations techniques. Key operational functions, including data collection and transmission, environment characterization, systems constraints, human factors, and navigation, often are intertwined and conflicted. Deep Reinforcement Learning (DRL) offers a framework for addressing integrated spacecraft navigation and planning in an uncertain dynamical environment. The goal of this study is to evaluate the utility of DRL for integrated spacecraft navigation and planning. This is achieved by developing a simple environmental characterization training environment in the Planar-Restricted 2-Body Problem (PR2BP), establishing benchmarks and heuristic baselines, and designing a previously unstudied Markov Decision Process (MDP) formulation. This MDP formulation enables the spacecraft DRL agents to appropriately balance navigation and actuation capabilities. The resulting DRL-derived policy exceeds a random or untrained policy and meets or exceeds the level of performance of a heuristic without actuation. In the process, valuable intuition is gained about the problem with insight into how DRL methods could scale to increasingly more realistic scenarios, including net-work design and training architectures, efficient state space representations, and methods for encouraging exploration in a parametric action space, among others.

navigation↗