Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
arXiv:2109.09628
Abstract
Self-supervised monocular depth prediction provides a cost-effective solution to obtain the 3D location of each pixel. However, the existing approaches usually lead to unsatisfactory accuracy, which is critical for autonomous robots. In this paper, we propose FusionDepth, a novel two-stage network to advance the self-supervised monocular dense depth learning by leveraging low-cost sparse (e.g. 4-beam) LiDAR. Unlike the existing methods that use sparse LiDAR mainly in a manner of time-consuming iterative post-processing, our model fuses monocular image features and sparse LiDAR features to predict initial depth maps. Then, an efficient feed-forward refine network is further designed to correct the errors in these initial depth maps in pseudo-3D space with real-time performance. Extensive experiments show that our proposed model significantly outperforms all the state-of-the-art self-supervised methods, as well as the sparse-LiDAR-based methods on both self-supervised monocular depth prediction and completion tasks. With the accurate dense depth prediction, our model outperforms the state-of-the-art sparse-LiDAR-based method (Pseudo-LiDAR++) by more than 68% for the downstream task monocular 3D object detection on the KITTI Leaderboard. Code is available at https://github.com/AutoAILab/FusionDepth
Accepted by CoRL2021
References in corpus (15)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Monocular 3D Object Detection and Box Fitting Trained End-to-End Using Intersection-over-Union Loss
- RTM3D: Real-time Monocular 3D Detection from Object Keypoints for Autonomous Driving
- MonoPair: Monocular 3D Object Detection Using Pairwise Spatial Relationships
- Feature-metric Loss for Self-supervised Learning of Depth and Egomotion
- CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion
- Learning monocular depth estimation infusing traditional stereo knowledge
- Dense Depth Posterior (DDP) from Single Image and Sparse Range
- Unsupervised High-Resolution Depth Learning From Videos With Dual Networks
- Rethinking Pseudo-LiDAR Representation
- End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection
- Robust Semi-Supervised Monocular Depth Estimation with Reprojected Distances
- Deep Learning based Monocular Depth Prediction: Datasets, Methods and Applications
- Parse Geometry from a Line: Monocular Depth Estimation with Partial Laser Observation
- Learning Joint 2D-3D Representations for Depth Completion