Veritatem Dies Aperit- Temporally Consistent Depth Prediction Enabled by a Multi-Task Geometric and Semantic Scene Understanding Approach
arXiv:1903.10764
Abstract
Robust geometric and semantic scene understanding is ever more important in many real-world applications such as autonomous driving and robotic navigation. In this paper, we propose a multi-task learning-based approach capable of jointly performing geometric and semantic scene understanding, namely depth prediction (monocular depth estimation and depth completion) and semantic scene segmentation. Within a single temporally constrained recurrent network, our approach uniquely takes advantage of a complex series of skip connections, adversarial training and the temporal constraint of sequential frame recurrence to produce consistent depth and semantic class labels simultaneously. Extensive experimental evaluation demonstrates the efficacy of our approach compared to other contemporary state-of-the-art techniques.
CVPR 2019
References in corpus (4)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Parallel Multi-Dimensional LSTM, With Application to Fast Biomedical Volumetric Image Segmentation
- STFCN: Spatio-Temporal FCN for Semantic Video Segmentation