Unsupervised Monocular Depth Estimation with Left-Right Consistency
arXiv:1609.03677
Abstract
Learning based methods have shown very promising results for the task of depth estimation in single images. However, most existing approaches treat depth prediction as a supervised regression problem and as a result, require vast quantities of corresponding ground truth depth data for training. Just recording quality depth data in a range of environments is a challenging problem. In this paper, we innovate beyond existing approaches, replacing the use of explicit depth data during training with easier-to-obtain binocular stereo footage. We propose a novel training objective that enables our convolutional neural network to learn to perform single image depth estimation, despite the absence of ground truth depth data. Exploiting epipolar geometry constraints, we generate disparity images by training our network with an image reconstruction loss. We show that solving for image reconstruction alone results in poor quality depth images. To overcome this problem, we propose a novel training loss that enforces consistency between the disparities produced relative to both the left and right images, leading to improved performance and robustness compared to existing approaches. Our method produces state of the art results for monocular depth estimation on the KITTI driving dataset, even outperforming supervised methods that have been trained with ground truth depth.
CVPR 2017 oral
References in corpus (13)
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- FlowNet: Learning Optical Flow with Convolutional Networks
- DepthTransfer: Depth Extraction from Video Using Non-parametric Sampling
- Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
- Virtual Worlds as Proxy for Multi-Object Tracking Analysis
- Spatio-temporal video autoencoder with differentiable memory
- Loss Functions for Neural Networks for Image Processing
- Single-Image Depth Perception in the Wild
- DeepStereo: Learning to Predict New Views from the World's Imagery
- Estimating Depth from Monocular Images as Classification Using Deep Fully Convolutional Residual Networks
- View Synthesis by Appearance Flow
- Learning the Matching Function
Cited by in corpus (14)
- SfM-Net: Learning of Structure and Motion from Video
- Sparse-to-Dense: Depth Prediction from Sparse Depth Samples and a Single Image
- SO-Net: Self-Organizing Network for Point Cloud Analysis
- Semi-Supervised Deep Learning for Monocular Depth Map Prediction
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
- Guided Optical Flow Learning
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations
- Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry
- Learning Single-Image Depth from Videos using Quality Assessment Networks
- Object Detection on Single Monocular Images through Canonical Correlation Analysis
- Monocular Depth Estimation with Directional Consistency by Deep Networks
- Deep Depth from Defocus: how can defocus blur improve 3D estimation using dense neural networks?
- Improving Self-Supervised Single View Depth Estimation by Masking Occlusion