SceneNet: Understanding Real World Indoor Scenes With Synthetic Data
arXiv:1511.07041
Abstract
Scene understanding is a prerequisite to many high level tasks for any automated intelligent machine operating in real world environments. Recent attempts with supervised learning have shown promise in this direction but also highlighted the need for enormous quantity of supervised data --- performance increases in proportion to the amount of data used. However, this quickly becomes prohibitive when considering the manual labour needed to collect such data. In this work, we focus our attention on depth based semantic per-pixel labelling as a scene understanding problem and show the potential of computer graphics to generate virtually unlimited labelled data from synthetic 3D scenes. By carefully synthesizing training data with appropriate noise models we show comparable performance to state-of-the-art RGBD systems on NYUv2 dataset despite using only depth data as input and set a benchmark on depth-based segmentation on SUN RGB-D dataset. Additionally, we offer a route to generating synthesized frame or video data, and understanding of different factors influencing performance gains.
References in corpus (1)
Cited by in corpus (28)
- The Replica Dataset: A Digital Replica of Indoor Spaces
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
- Spatio-temporal video autoencoder with differentiable memory
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- SemanticFusion: Dense 3D Semantic Mapping with Convolutional Neural Networks
- Object Detection Using Deep CNNs Trained on Synthetic Images
- Semantic Scene Completion from a Single Depth Image
- Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
- EdgeNet: Semantic Scene Completion from a Single RGB-D Image
- The AdobeIndoorNav Dataset: Towards Deep Reinforcement Learning based Real-world Indoor Robot Visual Navigation
- TorontoCity: Seeing the World with a Million Eyes
- Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras
- UnrealCV: Connecting Computer Vision to Unreal Engine
- Synthetic 3D Data Generation Pipeline for Geometric Deep Learning in Architecture
- DeepContext: Context-Encoding Neural Pathways for 3D Holistic Scene Understanding
- NVS-MonoDepth: Improving Monocular Depth Prediction with Novel View Synthesis
- Cut-and-Paste Dataset Generation for Balancing Domain Gaps in Object Instance Detection
- NViSII: A Scriptable Tool for Photorealistic Image Generation
- Fast Deep Matting for Portrait Animation on Mobile Phone
- BIM Hyperreality: Data Synthesis Using BIM and Hyperrealistic Rendering for Deep Learning
- 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation
- Pano2CAD: Room Layout From A Single Panorama Image
- Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks
- Learning Where to Look: Data-Driven Viewpoint Set Selection for 3D Scenes
- Monocular Spherical Depth Estimation with Explicitly Connected Weak Layout Cues
- 3D Scene Parsing via Class-Wise Adaptation
- Global and Local Texture Randomization for Synthetic-to-Real Semantic Segmentation