VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This report presents the results of a two-part project. The first part presents results of performance assessment tests on an Internet Library Information Assembly Data Base (ILIAD). It was found that ILLAD performed best when queries were short (one-to-three keywords), and were made up of rare, unambiguous words. In such cases as many as 64% of the typically 25 returned documents were found to be relevant. It was also found that a query format that was not so rigid with respect to spelling errors and punctuation marks would be more user-friendly. The second part of the report shows the design of a Kalman Filter for estimating motion parameters of a three dimensional object from sequences of noisy data derived from two-dimensional pictures. Given six measured deviation values represendng X, Y, Z, pitch, yaw, and roll, twelve parameters were estimated comprising the six deviations and their time rate of change. Values for the state transiton matrix, the observation matrix, the system noise covariance matrix, and the observation noise covariance matrix were determined. A simple way of initilizing the error covariance matrix was pointed out.
A navigation system includes an image acquisition device for acquiring a range image of a target vehicle, at least one processor, a memory including a target vehicle model and computer readable program code, where the processor and the computer readable program code are configured to cause the navigation system to convert the range image to a point cloud having three dimensions, compute a transform from the target vehicle model to the point cloud, and use the transform to estimate the target vehicle's attitude and position for capturing the target vehicle.
This data was created with open-source tooling for the purpose of running a public competition. This dataset is a sample of what can be generated for the commissioned competition runner to be able to set up their testing infrastructure. It contains images of publicly available spacecraft models in a Blender scene composed of a light source, a background image, and the spacecraft model with bounding box labels.
Explore the source record for details and available documents.
A system for determining the position, orientation and motion of a satellite with respect to a robotic spacecraft using video data is advanced. This system utilizes two levels of pose and motion estimation: an initial system which provides coarse estimates of pose and motion, and a second system which uses the coarse estimates and further processing to provide finer pose and motion estimates. The present paper emphasizes the initial coarse pose and motion estimation sybsystem. This subsystem utilizes novelty detection and filtering for locating novel parts and a neural net tracker to track these parts over time. Results of using this system on a sequence of images of a spin stabilized satellite are presented.
Future space missions require that spacecraft have the capability to autonomously navigate non-cooperative environments for rendezvous and proximity operations (RPO). Current relative navigation filters can have difficulty in these situations, diverging due to complications with data association, high measurement uncertainty, and clutter, particularly when detailed a priori maps of the target object or spacecraft do not exist. The goal of this work is to demonstrate the feasibility of random finite set (RFS) filters for spacecraft relative navigation and pose estimation. The approach is to formulate satellite relative navigation and pose estimation as a simultaneous localization and mapping (SLAM) problem, in which an observer spacecraft seeks to simultaneously estimate the location of features on a target object or spacecraft as well as its relative position, velocity and attitude. This work utilizes a filter developed using the framework of RFS which are well suited to multi-target SLAM operations, avoiding data association entirely. Relevant RPO scenarios with simulated flash LIDAR measurements are tested with a Probability Hypothesis Density (PHD) RFS filter embedded in a particle filter to obtain a feature map of a target and a relative pose estimate between the target and observer. Preliminary results show that an RFS-based filter can successfully perform SLAM in a spacecraft relative navigation scenario with no a priori map of the target. These results demonstrate the feasibility of RFS filtering for spacecraft relative navigation and motivate future studies which may expand to tracking space objects for space situational awareness, as well as relative navigation around small bodies.
Future space missions require that spacecraft have the capability to autonomously navigate non-cooperative environments for rendezvous and proximity operations (RPO). Current relative navigation filters can have difficulty in these situations, diverging due to complications with data association, high measurement uncertainty, and clutter, particularly when detailed a priori maps of the target object or spacecraft do not exist. The goal of this work is to demonstrate the feasibility of random finite set (RFS) filters for spacecraft relative navigation and pose estimation. The approach is to formulate satellite relative navigation and pose estimation as a simultaneous localization and mapping (SLAM) problem, in which an observer spacecraft seeks to simultaneously estimate the location of features on a target object or spacecraft as well as its relative position, velocity and attitude. This work utilizes a filter developed using the framework of RFS which are well suited to multi-target SLAM operations, avoiding data association entirely. Relevant RPO scenarios with simulated flash LIDAR measurements are tested with a Probability Hypothesis Density (PHD) RFS filter embedded in a particle filter to obtain a feature map of a target and a relative pose estimate between the target and observer. Preliminary results show that an RFS-based filter can successfully perform SLAM in a spacecraft relative navigation scenario with no a priori map of the target. These results demonstrate the feasibility of RFS filtering for spacecraft relative navigation and motivate future studies which may expand to tracking space objects for space situational awareness, as well as relative navigation around small bodies.
Monocular information from a gripper-mounted camera is used to servo the robot gripper to grasp a cylinder. The fundamental concept for rapid pose estimation is to reduce the amount of information that needs to be processed during each vision update interval. The grasping procedure is divided into four phases: learn, recognition, alignment, and approach. In the learn phase, a cylinder is placed in the gripper and the pose estimate is stored and later used as the servo target. This is performed once as a calibration step. The recognition phase verifies the presence of a cylinder in the camera field of view. An initial pose estimate is computed and uncluttered scan regions are selected. The radius of the cylinder is estimated by moving the robot a fixed distance toward the cylinder and observing the change in the image. The alignment phase processes only the scan regions obtained previously. Rapid pose estimates are used to align the robot with the cylinder at a fixed distance from it. The relative motion of the cylinder is used to generate an extrapolated pose-based trajectory for the robot controller. The approach phase guides the robot gripper to a grasping position. The cylinder can be grasped with a minimal reaction force and torque when only rough global pose information is initially available.
Monocular information from a gripper-mounted camera is used to servo the robot gripper to grasp a cylinder. The fundamental concept for rapid pose estimation is to reduce the amount of information that needs to be processed during each vision update interval. The grasping procedure is divided into four phases: learn, recognition, alignment, and approach. In the learn phase, a cylinder is placed in the gripper and the pose estimate is stored and later used as the servo target. This is performed once as a calibration step. The recognition phase verifies the presence of a cylinder in the camera field of view. An initial pose estimate is computed and uncluttered scan regions are selected. The radius of the cylinder is estimated by moving the robot a fixed distance toward the cylinder and observing the change in the image. The alignment phase processes only the scan regions obtained previously. Rapid pose estimates are used to align the robot with the cylinder at a fixed distance from it. The relative motion of the cylinder is used to generate an extrapolated pose-based trajectory for the robot controller. The approach phase guides the robot gripper to a grasping position. The cylinder can be grasped with a minimal reaction force and torque when only rough global pose information is initially available.
BACKGROUND Astronauts returning from long-duration exposure to microgravity frequently exhibit alterations in sensorimotor function leading to postural imbalance, impaired locomotion, and operational challenges to manual control. Mission duration and individual responses often influence both the severity of performance decrements and the variability in adaptation timelines. Postflight disruptions during functional tasks are often detected through body-worn inertial measurement unit (IMU) devices. While IMU sensors are relatively compact, the long-term wear may lead to discomfort, displacement of the sensors on the body, and restrictions in movement or crew behavior. Although IMU data offers valuable insights from a research standpoint, interpreting changes in pre- and post-flight measures can be difficult for crew support personnel beyond the research domain, which can hinder the application for medical assessments and rehabilitation. Finally, the availability of inertial sensors in-flight is limited. There is a need for unobtrusive monitoring tools to improve our ability to monitor adaptation following gravitational transitions in various postflight evaluations and rehabilitation settings. Markerless motion capture (MMC) is an evolving unobtrusive technology that builds upon decades of research with marker-based motion capture systems to provide 3D human pose estimation from multiple synchronized 2D camera views using deep learning algorithms. Markerless technology can revolutionize how data is captured pre- and post-flight and potentially in-flight during intravehicular activity by enabling pose estimation of multiple crew members from onboard camera hardware. METHODS The following presents the initial evaluation of a state-of-the-art commercial-off-the-shelf MMC system, Theia Markerless, compared to IMU devices during various ground-based functional tasks and environmental conditions. The featured functional tasks include assessments from Human Research Program (HRP) funded studies such as Sensorimotor Standard Measures and Sensorimotor Assessments. Synchronous data collected using both motion capture and IMUs are analyzed for six male and female subjects of varying anthropometry. The analysis includes limited assessments of clothing, capture volume configurations, and the tool's sensitivity to detecting performance changes after a spaceflight analog centrifuge exposure. The development of visualization tools to enhance the application of the pose estimation output is also presented. RESULTS Initial results demonstrate comparable root mean square error (RMSE) to existing literature evaluating markerless and marker-based motion capture systems. Considering the relative functional range of motion of the cervical spine, normalized error values for the markerless system’s accuracy of the head was 0.032 in pitch, 0.025 in roll, and 0.018 in yaw plane of motion across a subset of functional tasks. The raw RMSE values were 3.49, 2.25, and 2.81 degrees respectively. For the torso, results suggest normalized errors of 0.469 in pitch, 0.191 in roll, and 0.307 in yaw planes of motion and raw RMSE values of 3.52, 1.53, and 3.07 degrees respectively. The data suggests the functional demands of a particular task influences the estimation accuracy of the MMC system where more dynamic motion and cases where subjects are not upright may introduce diminished tracking accuracy. DISCUSSION The following work lays the foundation for future implementations leveraging markerless motion capture to assess the time course of recovery and provide insight for rehabilitation protocols to enhance crew readiness for the resumption of daily activities. These tools offer effective methods for anonymizing sensitive crew data, facilitating numerous applications across research, medical, and rehabilitation groups. Collaborations with the Anthropometry and Biomechanics Facility will provide further comparisons of the Markerless system to a marker-based system. ACKNOWLEDGEMENT</ This work is supported by NASA’s Exploration Systems Development Mission Directorate Mars Campaign Office Crew Health Countermeasures.
In-space assembly operations require accurate reasoning over the pose, location, and structural organization of both the autonomous agents and assembly materials. In a full six-degree-of-freedom space, an accurate understanding of the full three-dimensional structure of the object of interest greatly enriches information for pose estimation and collision planning. Current methods of predicting pose estimation require a priori understanding of the shape of the object. Additionally, visual information in the space environment is impacted by variations in contrast and illumination. Using synthetic data allows us to rapidly generate large datasets with in varying environments and lighting conditions.This work details the generation of synthetic data used to explore the use of a region-based convolutional neural networks to detect objects of interest and predict a voxel-based three-dimensional mesh in order to understand their full three-dimensional shape. This mesh provides useful spatial information during in-space assembly operations without requiring either the complexity of maintaining models over the progress of building an object or observations from multiple angles. The generated meshes are then compared to that of ground truth in order to measure its performance.
In-space assembly operations require accurate reasoning over the pose, location, and structural organization of both the autonomous agents and assembly materials. In a full six-degree-of-freedom space, an accurate understanding of the full three-dimensional structure of the object of interest greatly enriches information for pose estimation and collision planning. Current methods of predicting pose estimation require a priori understanding of the shape of the object. Additionally, visual information in the space environment is impacted by variations in contrast and illumination. Using synthetic data allows us to rapidly generate large datasets with in varying environments and lighting conditions. This work details the generation of synthetic data used to explore the use of a region-based convolutional neural networks to detect objects of interest and predict a voxel-based three-dimensional mesh in order to understand their full three-dimensional shape. This mesh provides useful spatial information during in-space assembly operations without requiring either the complexity of maintaining models over the progress of building an object or observations from multiple angles. The generated meshes are then compared to that of ground truth in order to measure its performance.
The motivation for vision-guided servoing is taken from tasks in automated or telerobotic space assembly and construction. Vision-guided servoing requires the ability to perform rapid pose estimates and provide predictive feature tracking. Monocular information from a gripper-mounted camera is used to servo the gripper to grasp a cylinder. The procedure is divided into recognition and servo phases. The recognition stage verifies the presence of a cylinder in the camera field of view. Then an initial pose estimate is computed and uncluttered scan regions are selected. The servo phase processes only the selected scan regions of the image. Given the knowledge, from the recognition phase, that there is a cylinder in the image and knowing the radius of the cylinder, 4 of the 6 pose parameters can be estimated with minimal computation. The relative motion of the cylinder is obtained by using the current pose and prior pose estimates. The motion information is then used to generate a predictive feature-based trajectory for the path of the gripper.
The relational pyramid representation is a hierarchical relational description of an object that can be used in recognition and pose estimation. In this representation, primitive features appear at the bottom of the pyramid, relations among primitives appear at level one, and, in general, relations among lower-level structures appear at the levels above the levels in which these structures were defined. A pose-estimation system is constructed that uses the relational pyramid representation for the view classes of a 3D object and for the description of the unknown view of an object extracted from an image. A previously described system uses summary information to rank and select view classes to which the unknown view is compared. This paper describes a best-first search procedure to find correspondences between the selected view class pyramid and the image pyramid.
Future space missions require spacecraft to autonomously navigate non-cooperative environments for rendezvous and proximity operations (RPO). Current navigation filters used for RPO can have difficulty when optical sensors are used, due to complications with data association, high measurement uncertainty, and clutter. This paper provides an initial demonstration of the feasibility of using random finite set (RFS) filters for spacecraft relative navigation and pose estimation. Spacecraft relative navigation is formulated as a simultaneous localization and mapping (SLAM) problem, in which an observer spacecraft seeks to simultaneously estimate the location of features on a target object or spacecraft as well as its relative position, velocity and attitude. Simulated flash LIDAR measurements are processed using a Gaussian Mixture Probability Hypothesis Density (GMPHD) filter embedded in a particle filter to obtain a feature map of a target and a relative pose estimate between the target and observer over time. Results show that an RFS-based filter can successfully perform SLAM in a spacecraft relative navigation scenario.
Research is being conducted at the LaRC to develop a telerobotic assembly system designed to construct large space truss structures. This research program was initiated within the past several years, and a ground-based test-bed was developed to evaluate and expand the state of the art. Test-bed operations currently use predetermined ('taught') points for truss structural assembly. Total dependence on the use of taught points for joint receptacle capture and strut installation is neither robust nor reliable enough for space operations. Therefore, a machine vision sensor guidance system is being developed to locate and guide the robot to a passive target mounted on the truss joint receptacle. The vision system hardware includes a miniature video camera, passive targets mounted on the joint receptacles, target illumination hardware, and an image processing system. Discrimination of the target from background clutter is accomplished through standard digital processing techniques. Once the target is identified, a pose estimation algorithm is invoked to determine the location, in three-dimensional space, of the target relative to the robots end-effector. Preliminary test results of the vision system in the Automated Structural Assembly Laboratory with a range of lighting and background conditions indicate that it is fully capable of successfully identifying joint receptacle targets throughout the required operational range. Controlled optical bench test results indicate that the system can also provide the pose estimation accuracy to define the target position.
This paper describes a machine vision system for relative spacecraft navigation during the terminal phase of approach to docking that: 1) matches high contrast image features of the target vehicle, as seen by a camera that is bore-sighted to the docking adapter on the chase vehicle, to the corresponding features in a 3d model of the docking adapter on the target vehicle and 2) is robust to on-orbit lighting. An implementation is provided for the case of the Space Shuttle Orbiter docking to the International Space Station (ISS) with quantitative test results using a full scale, medium fidelity mock-up of the ISS docking adapter mounted on a 6-DOF motion platform at the NASA Marshall Spaceflight Center Flight Robotics Laboratory and qualitative test results using recorded video from the Orbiter Docking System Camera (ODSC) during multiple orbiter to ISS docking missions. The Natural Feature Image Registration (NFIR) system consists of two modules: 1) Tracking which tracks the target object from image to image and estimates the position and orientation (pose) of the docking camera relative to the target object and 2) Acquisition which recognizes the target object if it is in the docking camera Field-of-View and provides an approximate pose that is used to initialize tracking. Detected image edges are matched to the 3d model edges whose predicted location, based on the pose estimate and its first time derivative from the previous frame, is closest to the detected edge1 . Mismatches are eliminated using a rigid motion constraint. The remaining 2d image to 3d model matches are used to make a least squares estimate of the change in relative pose from the previous image to the current image. The changes in position and in attitude are used as data for two Kalman filters whose outputs are smoothed estimate of position and velocity plus attitude and attitude rate that are then used to predict the location of the 3d model features in the next image.