Calibrating Self-supervised Monocular Depth Estimation
arXiv:2009.07714
Abstract
In the recent years, many methods demonstrated the ability of neural networks to learn depth and pose changes in a sequence of images, using only self-supervision as the training signal. Whilst the networks achieve good performance, the often over-looked detail is that due to the inherent ambiguity of monocular vision they predict depth up to an unknown scaling factor. The scaling factor is then typically obtained from the LiDAR ground truth at test time, which severely limits practical applications of these methods. In this paper, we show that incorporating prior information about the camera configuration and the environment, we can remove the scale ambiguity and predict depth directly, still using the self-supervised formulation and not relying on any additional sensors.
References in corpus (7)
- From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
- Deep Ordinal Regression Network for Monocular Depth Estimation
- Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency
- Learning Monocular Depth by Distilling Cross-domain Stereo Networks
- Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding
- Monocular Depth Estimation with Self-supervised Instance Adaptation
- Single View Stereo Matching