CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
arXiv:1704.03489
Abstract
Given the recent advances in depth prediction from Convolutional Neural Networks (CNNs), this paper investigates how predicted depth maps from a deep neural network can be deployed for accurate and dense monocular reconstruction. We propose a method where CNN-predicted dense depth maps are naturally fused together with depth measurements obtained from direct monocular SLAM. Our fusion scheme privileges depth prediction in image locations where monocular SLAM approaches tend to fail, e.g. along low-textured regions, and vice-versa. We demonstrate the use of depth prediction for estimating the absolute scale of the reconstruction, hence overcoming one of the major limitations of monocular SLAM. Finally, we propose a framework to efficiently fuse semantic labels, obtained from a single frame, with dense SLAM, yielding semantically coherent scene reconstruction from a single view. Evaluation results on two benchmark datasets show the robustness and accuracy of our approach.
10 pages, 6 figures, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), Hawaii, USA, June, 2017. The first two authors contribute equally to this paper
Cited by in corpus (25)
- Learning Depth from Monocular Videos using Direct Methods
- RSGM: Real-time Raster-Respecting Semi-Global Matching for Power-Constrained Systems
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
- FutureMapping: The Computational Structure of Spatial AI Systems
- Automatic Generation of Constrained Furniture Layouts
- RegNet: Learning the Optimization of Direct Image-to-Image Pose Registration
- Visual Odometry Revisited: What Should Be Learnt?
- Self-Supervised Learning of Depth and Camera Motion from 360° Videos
- ENG: End-to-end Neural Geometry for Robust Depth and Pose Estimation using CNNs
- Recurrent Neural Network for Learning DenseDepth and Ego-Motion from Video
- Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry
- DeLS-3D: Deep Localization and Segmentation with a 3D Semantic Map
- WGANVO: Monocular Visual Odometry based on Generative Adversarial Networks
- 3D-GMNet: Single-View 3D Shape Recovery as A Gaussian Mixture
- Dense RGB-D semantic mapping with Pixel-Voxel neural network
- On Machine Learning and Structure for Mobile Robots
- Semantic Photometric Bundle Adjustment on Natural Sequences
- DeepTAM: Deep Tracking and Mapping
- Unsupervised Learning-based Depth Estimation aided Visual SLAM Approach
- CodeMapping: Real-Time Dense Mapping for Sparse SLAM using Compact Scene Representations
- Self-Supervised Monocular Image Depth Learning and Confidence Estimation
- Position Estimation of Camera Based on Unsupervised Learning
- Direct Visual-Inertial Odometry with Semi-Dense Mapping
- Incremental Scene Synthesis
- Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space