RotationNet: Joint Object Categorization and Pose Estimation Using Multiviews from Unsupervised Viewpoints
arXiv:1603.06208
Abstract
We propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels for training, our method treats the viewpoint labels as latent variables, which are learned in an unsupervised manner during the training using an unaligned object dataset. RotationNet is designed to use only a partial set of multi-view images for inference, and this property makes it useful in practical scenarios where only partial views are available. Moreover, our pose alignment strategy enables one to obtain view-specific feature representations shared across classes, which is important to maintain high accuracy in both object categorization and pose estimation. Effectiveness of RotationNet is demonstrated by its superior performance to the state-of-the-art methods of 3D object classification on 10- and 40-class ModelNet datasets. We also show that RotationNet, even trained without known poses, achieves the state-of-the-art performance on an object pose estimation dataset. The code is available on https://github.com/kanezaki/rotationnet
24 pages, 23 figures. Accepted to CVPR 2018
References in corpus (8)
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Generative and Discriminative Voxel Modeling with Convolutional Neural Networks
- FPNN: Field Probing Neural Networks for 3D Data
- Unsupervised Learning of Depth and Ego-Motion from Video
- FusionNet: 3D Object Classification Using Multiple Data Representations
- GIFT: A Real-time and Scalable 3D Shape Search Engine
- Semantic Pose using Deep Networks Trained on Synthetic RGB-D
Cited by in corpus (6)
- A Survey on Deep Learning Methods for Robot Vision
- Point2Sequence: Learning the Shape Representation of 3D Point Clouds with an Attention-based Sequence to Sequence Network
- Volumetric Convolution: Automatic Representation Learning in Unit Ball
- Method for the generation of depth images for view-based shape retrieval of 3D CAD model from partial point cloud
- A Variational Feature Encoding Method of 3D Object for Probabilistic Semantic SLAM
- A Deeper Look at 3D Shape Classifiers