Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras
arXiv:1703.08866
Abstract
Visual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel approach to object-class segmentation from multiple RGB-D views using deep learning. We train a deep neural network to predict object-class semantics that is consistent from several view points in a semi-supervised way. At test time, the semantics predictions of our network can be fused more consistently in semantic keyframe maps than predictions of a network trained on individual views. We base our network architecture on a recent single-view deep learning approach to RGB and depth fusion for semantic object-class segmentation and enhance it with multi-scale loss minimization. We obtain the camera trajectory using RGB-D SLAM and warp the predictions of RGB-D images into ground-truth annotated frames in order to enforce multi-view consistency during training. At test time, predictions from multiple views are fused into keyframes. We propose and analyze several methods for enforcing multi-view consistency during training and testing. We evaluate the benefit of multi-view consistency training and demonstrate that pooling of deep features and fusion over multiple views outperforms single-view baselines on the NYUDv2 benchmark for semantic segmentation. Our end-to-end trained network achieves state-of-the-art performance on the NYUDv2 dataset in single-view segmentation as well as multi-view semantic fusion.
the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2017)
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Learning Deconvolution Network for Semantic Segmentation
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Indoor Semantic Segmentation using depth information
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- SceneNet: Understanding Real World Indoor Scenes With Synthetic Data
- SemanticFusion: Dense 3D Semantic Mapping with Convolutional Neural Networks
- OctNet: Learning Deep 3D Representations at High Resolutions
- Exploring Context with Deep Structured models for Semantic Segmentation
Cited by in corpus (7)
- A Review on Deep Learning Techniques Applied to Semantic Segmentation
- Sparse Bayesian Inference for Dense Semantic Mapping
- Dense RGB-D semantic mapping with Pixel-Voxel neural network
- Recurrent-OctoMap: Learning State-based Map Refinement for Long-Term Semantic Mapping with 3D-Lidar Data
- SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
- In pixels we trust: From Pixel Labeling to Object Localization and Scene Categorization
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views