Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
arXiv:2010.04030 · doi:10.1016/j.cviu.2022.103440
Abstract
Representing scenes at the granularity of objects is a prerequisite for scene understanding and decision making. We propose PriSMONet, a novel approach based on Prior Shape knowledge for learning Multi-Object 3D scene decomposition and representations from single images. Our approach learns to decompose images of synthetic scenes with multiple objects on a planar surface into its constituent scene objects and to infer their 3D properties from a single view. A recurrent encoder regresses a latent representation of 3D shape, pose and texture of each object from an input RGB image. By differentiable rendering, we train our model to decompose scenes from RGB-D images in a self-supervised way. The 3D shapes are represented continuously in function-space as signed distance functions which we pre-train from example shapes in a supervised way. These shape priors provide weak supervision signals to better condition the challenging overall learning task. We evaluate the accuracy of our model in inferring 3D scene layout, demonstrate its generative capabilities, assess its generalization to real images, and point out benefits of the learned representation.
Preprint accepted to Computer Vision and Image Understanding
References in corpus (13)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- MarrNet: 3D Shape Reconstruction via 2.5D Sketches
- Object-Centric Learning with Slot Attention
- Learning a Multi-View Stereo Machine
- Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction
- Differentiable Rendering: A Survey
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- Decomposing 3D Scenes into Objects via Unsupervised Volume Segmentation
- Learning Object-Centric Representations of Multi-Object Scenes from Multiple Views
- ROOTS: Object-Centric Representation and Rendering of 3D Scenes
- Unsupervised object-centric video generation and decomposition in 3D