Semantically-Guided Representation Learning for Self-Supervised Monocular Depth
arXiv:2002.12319
Abstract
Self-supervised learning is showing great promise for monocular depth estimation, using geometry as the only source of supervision. Depth networks are indeed capable of learning representations that relate visual appearance to 3D properties by implicitly leveraging category-level patterns. In this work we investigate how to leverage more directly this semantic structure to guide geometric representation learning, while remaining in the self-supervised regime. Instead of using semantic labels and proxy losses in a multi-task approach, we propose a new architecture leveraging fixed pretrained semantic segmentation networks to guide self-supervised representation learning via pixel-adaptive convolutions. Furthermore, we propose a two-stage training process to overcome a common semantic bias on dynamic objects via resampling. Our method improves upon the state of the art for self-supervised monocular depth prediction over all pixels, fine-grained details, and per semantic categories.
Proceedings of the Eighth International Conference on Learning Representations (ICLR 2020)
References in corpus (2)
Cited by in corpus (6)
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Forget About the LiDAR: Self-Supervised Depth Estimators with MED Probability Volumes
- PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single-View Depth Estimation with Neural Positional Encoding and Distilled Matting Loss
- Semantic-Guided Representation Enhancement for Self-supervised Monocular Trained Depth Estimation
- CI-Net: Contextual Information for Joint Semantic Segmentation and Depth Estimation
- Domain Adaptive Monocular Depth Estimation With Semantic Information