Unsupervised Learning of 3D Structure from Images
arXiv:1607.00662
Abstract
A key goal of computer vision is to recover the underlying 3D structure from 2D observations of the world. In this paper we learn strong deep generative models of 3D structures, and recover these structures from 3D and 2D images via probabilistic inference. We demonstrate high-quality samples and report log-likelihoods on several datasets, including ShapeNet [2], and establish the first benchmarks in the literature. We also show how these models and their inference networks can be trained end-to-end from 2D images. This demonstrates for the first time the feasibility of learning to infer 3D representations of the world in a purely unsupervised manner.
Appears in Advances in Neural Information Processing Systems 29 (NIPS 2016)
References in corpus (4)
Cited by in corpus (35)
- Unsupervised Learning of Depth and Ego-Motion from Video
- Learning a Multi-View Stereo Machine
- Deep Active Inference
- ASMFS: Adaptive-Similarity-based Multi-modality Feature Selection for Classification of Alzheimer's Disease
- Improved Adversarial Systems for 3D Object Generation and Reconstruction
- GIFS: Neural Implicit Function for General Shape Representation
- Deep Level Sets: Implicit Surface Representations for 3D Shape Inference
- A Point Set Generation Network for 3D Object Reconstruction from a Single Image
- Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling
- Multi-View Image Generation from a Single-View
- Hierarchical Surface Prediction for 3D Object Reconstruction
- SurfNet: Generating 3D shape surfaces using deep residual networks
- Compact Model Representation for 3D Reconstruction
- Octree Generating Networks: Efficient Convolutional Architectures for High-resolution 3D Outputs
- Interactive 3D Modeling with a Generative Adversarial Network
- 3D Shape Induction from 2D Views of Multiple Objects
- Unsupervised Learning of Shape and Pose with Differentiable Point Clouds
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- RGBD-GAN: Unsupervised 3D Representation Learning From Natural Image Datasets via RGBD Image Synthesis
- Learning Shape Priors for Single-View 3D Completion and Reconstruction
- Leveraging the Exact Likelihood of Deep Latent Variable Models
- View Inter-Prediction GAN: Unsupervised Representation Learning for 3D Shapes by Learning Global Shape Memories to Support Local View Predictions
- Deep Supervision with Shape Concepts for Occlusion-Aware 3D Object Parsing
- Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations
- Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
- VConv-DAE: Deep Volumetric Shape Learning Without Object Labels
- Layer-structured 3D Scene Inference via View Synthesis
- Dual-Camera All-in-Focus Neural Radiance Fields
- Active Perception and Representation for Robotic Manipulation
- Following Gaze Across Views
- 3D Shape Reconstruction from a Single 2D Image via 2D-3D Self-Consistency
- Learning a Hierarchical Latent-Variable Model of 3D Shapes
- Self-supervised 3D Shape and Viewpoint Estimation from Single Images for Robotics
- Deep disentangled representations for volumetric reconstruction
- iSPA-Net: Iterative Semantic Pose Alignment Network