Pix2Vox: Context-aware 3D Reconstruction from Single and Multi-view Images
arXiv:1901.11153 · doi:10.1109/ICCV.2019.00278
Abstract
Recovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method.
ICCV 2019
References in corpus (2)
Cited by in corpus (15)
- Artificial Intelligence in the Creative Industries: A Review
- Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
- Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
- PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction
- Implicit Full Waveform Inversion with Deep Neural Representation
- CSDN: Cross-modal Shape-transfer Dual-refinement Network for Point Cloud Completion
- DoF-NeRF: Depth-of-Field Meets Neural Radiance Fields
- SoftPool++: An Encoder-Decoder Network for Point Cloud Completion
- Estimating and abstracting the 3D structure of bones using neural networks on X-ray (2D) images
- Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
- 3D Teeth Reconstruction from Panoramic Radiographs using Neural Implicit Functions
- Dual-Camera All-in-Focus Neural Radiance Fields
- Mesh deformation-based single-view 3D reconstruction of thin eyeglasses frames with differentiable rendering
- Shape2.5D: A Dataset of Texture-less Surfaces for Depth and Normals Estimation
- View-Adaptive Renderer for View-Consistent 2D-to-3D Generation