Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images
arXiv:1508.04546
Abstract
Analysis-by-synthesis has been a successful approach for many tasks in computer vision, such as 6D pose estimation of an object in an RGB-D image which is the topic of this work. The idea is to compare the observation with the output of a forward process, such as a rendered image of the object of interest in a particular pose. Due to occlusion or complicated sensor noise, it can be difficult to perform this comparison in a meaningful way. We propose an approach that "learns to compare", while taking these difficulties into account. This is done by describing the posterior density of a particular object pose with a convolutional neural network (CNN) that compares an observed and rendered image. The network is trained with the maximum likelihood paradigm. We observe empirically that the CNN does not specialize to the geometry or appearance of specific objects, and it can be used with objects of vastly different shapes and appearances, and in different backgrounds. Compared to state-of-the-art, we demonstrate a significant improvement on two different datasets which include a total of eleven objects, cluttered background, and heavy occlusion.
16 pages, 8 figures
References in corpus (4)
Cited by in corpus (16)
- PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes
- Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image
- Siamese Regression Networks with Efficient mid-level Feature Extraction for 3D Object Pose Estimation
- BOP: Benchmark for 6D Object Pose Estimation
- CPS++: Improving Class-level 6D Pose and Shape Estimation From Monocular Images With Self-Supervised Learning
- A Deep Learning Approach for Pose Estimation from Volumetric OCT Data
- When Regression Meets Manifold Learning for Object Recognition and Pose Estimation
- Multi-Task Deep Networks for Depth-Based 6D Object Pose and Joint Registration in Crowd Scenarios
- A Unified Framework for Multi-View Multi-Class Object Pose Estimation
- DSAC - Differentiable RANSAC for Camera Localization
- Global Hypothesis Generation for 6D Object Pose Estimation
- Deep Gated Multi-modal Learning: In-hand Object Pose Changes Estimation using Tactile and Image Data
- PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning
- 6D Object Pose Estimation Based on 2D Bounding Box
- Recovering 6D Object Pose: A Review and Multi-modal Analysis
- REST: Real-to-Synthetic Transform for Illumination Invariant Camera Localization