Semi-Supervised Deep Learning for Monocular Depth Map Prediction
arXiv:1702.02706
Abstract
Supervised deep learning often suffers from the lack of sufficient training data. Specifically in the context of monocular depth map prediction, it is barely possible to determine dense ground truth depth images in realistic dynamic outdoor environments. When using LiDAR sensors, for instance, noise is present in the distance measurements, the calibration between sensors cannot be perfect, and the measurements are typically much sparser than the camera images. In this paper, we propose a novel approach to depth map prediction from monocular images that learns in a semi-supervised way. While we use sparse ground-truth depth for supervised learning, we also enforce our deep network to produce photoconsistent dense depth maps in a stereo setup using a direct image alignment loss. In experiments we demonstrate superior performance in depth map prediction from single images compared to the state-of-the-art methods.
CVPR 2017 Spotlight
References in corpus (2)
Cited by in corpus (40)
- Unsupervised Learning of Depth and Ego-Motion from Video
- DeepV2D: Video to Depth with Differentiable Structure from Motion
- Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency
- Sparse-to-Dense: Depth Prediction from Sparse Depth Samples and a Single Image
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Cascade Residual Learning: A Two-stage Convolutional Neural Network for Stereo Matching
- Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
- Learning Depth with Convolutional Spatial Propagation Network
- Monocular Depth Estimation using Multi-Scale Continuous CRFs as Sequential Deep Networks
- Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding
- Camera-based vehicle velocity estimation from monocular video
- LEGO: Learning Edge with Geometry all at Once by Watching Videos
- Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps with Accurate Object Boundaries
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- Sparse and Dense Data with CNNs: Depth Completion and Semantic Segmentation
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- Single View Stereo Matching
- Geometry meets semantics for semi-supervised monocular depth estimation
- Left-Right Comparative Recurrent Model for Stereo Matching
- T2Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks
- ENG: End-to-end Neural Geometry for Robust Depth and Pose Estimation using CNNs
- Every Pixel Counts: Unsupervised Geometry Learning with Holistic 3D Motion Understanding
- ActiveStereoNet: End-to-End Self-Supervised Learning for Active Stereo Systems
- Recurrent Neural Network for Learning DenseDepth and Ego-Motion from Video
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry
- SegStereo: Exploiting Semantic Information for Disparity Estimation
- Zoom and Learn: Generalizing Deep Stereo Matching to Novel Domains
- DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency
- UnDEMoN 2.0: Improved Depth and Ego Motion Estimation through Deep Image Sampling
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- Control of the Final-Phase of Closed-Loop Visual Grasping using Image-Based Visual Servoing
- Self-Supervised Monocular Image Depth Learning and Confidence Estimation
- Learning Single-Image Depth from Videos using Quality Assessment Networks
- Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- DAN: A Deformation-Aware Network for Consecutive Biomedical Image Interpolation
- Semantics-Driven Unsupervised Learning for Monocular Depth and Ego-Motion Estimation
- Balanced Depth Completion between Dense Depth Inference and Sparse Range Measurements via KISS-GP
- Unsupervised monocular stereo matching