Predicting Complete 3D Models of Indoor Scenes
arXiv:1504.02437
Abstract
One major goal of vision is to infer physical models of objects, surfaces, and their layout from sensors. In this paper, we aim to interpret indoor scenes from one RGBD image. Our representation encodes the layout of walls, which must conform to a Manhattan structure but is otherwise flexible, and the layout and extent of objects, modeled with CAD-like 3D shapes. We represent both the visible and occluded portions of the scene, producing a complete 3D parse. Such a scene interpretation is useful for robotics and visual reasoning, but difficult to produce due to the well-known challenge of segmentation, the high degree of occlusion, and the diversity of objects in indoor scene. We take a data-driven approach, generating sets of potential object regions, matching to regions in training images, and transferring and aligning associated 3D models while encouraging fit to observations and overall consistency. We demonstrate encouraging results on the NYU v2 dataset and highlight a variety of interesting directions for future work.
References in corpus (2)
Cited by in corpus (15)
- Semantic Scene Completion from a Single Depth Image
- LayoutNet: Reconstructing the 3D Room Layout from a Single RGB Image
- EdgeNet: Semantic Scene Completion from a Single RGB-D Image
- Semantic Scene Completion Combining Colour and Depth: preliminary experiments
- Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene Completion
- Deep Reinforcement Learning of Volume-guided Progressive View Inpainting for 3D Point Scene Completion from a Single Depth Image
- View-volume Network for Semantic Scene Completion from a Single Depth Image
- Efficient Semantic Scene Completion Network with Spatial Group Convolution
- Learning Direct Optimization for Scene Understanding
- GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes
- IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement
- Depth Based Semantic Scene Completion with Position Importance Aware Loss
- Complete 3D Scene Parsing from an RGBD Image
- OmniLayout: Room Layout Reconstruction from Indoor Spherical Panoramas
- Learning to Parse Wireframes in Images of Man-Made Environments