Learning Descriptors for Object Recognition and 3D Pose Estimation
arXiv:1502.05908 · doi:10.1109/CVPR.2015.7298930
Abstract
Detecting poorly textured objects and estimating their 3D pose reliably is still a very challenging problem. We introduce a simple but powerful approach to computing descriptors for object views that efficiently capture both the object identity and 3D pose. By contrast with previous manifold-based approaches, we can rely on the Euclidean distance to evaluate the similarity between descriptors, and therefore use scalable Nearest Neighbor search methods to efficiently handle a large number of objects under a large range of poses. To achieve this, we train a Convolutional Neural Network to compute these descriptors by enforcing simple similarity and dissimilarity constraints between the descriptors. We show that our constraints nicely untangle the images from different objects and different views into clusters that are not only well-separated but also structured as the corresponding sets of poses: The Euclidean distance between descriptors is large when the descriptors are from different objects, and directly related to the distance between the poses when the descriptors are from the same object. These important properties allow us to outperform state-of-the-art object views representations on challenging RGB and RGB-D data.
CVPR 2015
Cited by in corpus (28)
- SegMap: Segment-based mapping and localization using data-driven descriptors
- SegMap: 3D Segment Mapping using Data-Driven Descriptors
- Going Further with Point Pair Features
- Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
- Deep 6-DOF Tracking
- Rotational Subgroup Voting and Pose Clustering for Robust 3D Object Recognition
- Learning Correspondence Structures for Person Re-identification
- Occlusion-Aware Self-Supervised Monocular 6D Object Pose Estimation
- DPODv2: Dense Correspondence-Based 6 DoF Pose Estimation
- Learning to See the Wood for the Trees: Deep Laser Localization in Urban and Natural Environments on a CPU
- 3D Object Instance Recognition and Pose Estimation Using Triplet Loss with Dynamic Margin
- MinkLoc3D-SI: 3D LiDAR place recognition with sparse convolutions, spherical coordinates, and intensity
- 6D Pose Estimation with Combined Deep Learning and 3D Vision Techniques for a Fast and Accurate Object Grasping
- Domain-invariant Similarity Activation Map Contrastive Learning for Retrieval-based Long-term Visual Localization
- Real-Time Object Pose Estimation with Pose Interpreter Networks
- When Regression Meets Manifold Learning for Object Recognition and Pose Estimation
- 3D Object Detection and Pose Estimation of Unseen Objects in Color Images with Local Surface Embeddings
- GDRNPP: A Geometry-guided and Fully Learning-based Object Pose Estimator
- Real-Time 6D Object Pose Estimation on CPU
- Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
- Learning Social Image Embedding with Deep Multimodal Attention Networks
- View-Invariant, Occlusion-Robust Probabilistic Embedding for Human Pose
- Learning Orientation Distributions for Object Pose Estimation
- LCD -- Line Clustering and Description for Place Recognition
- DeepHMap++: Combined Projection Grouping and Correspondence Learning for Full DoF Pose Estimation
- Generative Model with Coordinate Metric Learning for Object Recognition Based on 3D Models
- A Small Form Factor Aerial Research Vehicle for Pick-and-Place Tasks with Onboard Real-Time Object Detection and Visual Odometry
- Self-supervised Latent Space Optimization with Nebula Variational Coding