Inferring Point Clouds from Single Monocular Images by Depth Intermediation
arXiv:1812.01402
Abstract
In this paper, we propose a pipeline to generate 3D point cloud of an object from a single-view RGB image. Most previous work predict the 3D point coordinates from single RGB images directly. We decompose this problem into depth estimation from single images and point cloud completion from partial point clouds. Our method sequentially predicts the depth maps from images and then infers the complete 3D object point clouds based on the predicted partial point clouds. We explicitly impose the camera model geometrical constraint in our pipeline and enforce the alignment of the generated point clouds and estimated depth maps. Experimental results for the single image 3D object reconstruction task show that the proposed method outperforms existing state-of-the-art methods. Both the qualitative and quantitative results demonstrate the generality and suitability of our method.
Statement: This paper is under consideration at Computer Vision and Image Understanding
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Generative and Discriminative Voxel Modeling with Convolutional Neural Networks
- Learning to Reconstruct Shapes from Unseen Classes
- Mining Point Cloud Local Structures by Kernel Correlation and Graph Pooling
- 3D-LMNet: Latent Embedding Matching for Accurate and Diverse 3D Point Cloud Reconstruction from a Single Image
- 3D Object Reconstruction from a Single Depth View with Adversarial Learning
- Three for one and one for three: Flow, Segmentation, and Surface Normals
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
Cited by in corpus (4)
- Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
- SE-MD: A Single-encoder multiple-decoder deep network for point cloud generation from 2D images
- Meta Deformation Network: Meta Functionals for Shape Correspondence
- Learnable Triangulation for Deep Learning-based 3D Reconstruction of Objects of Arbitrary Topology from Single RGB Images