MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer
arXiv:2208.03543 · doi:10.1109/3DV57658.2022.00077
Abstract
Self-supervised monocular depth estimation is an attractive solution that does not require hard-to-source depth labels for training. Convolutional neural networks (CNNs) have recently achieved great success in this task. However, their limited receptive field constrains existing network architectures to reason only locally, dampening the effectiveness of the self-supervised paradigm. In the light of the recent successes achieved by Vision Transformers (ViTs), we propose MonoViT, a brand-new framework combining the global reasoning enabled by ViT models with the flexibility of self-supervised monocular depth estimation. By combining plain convolutions with Transformer blocks, our model can reason locally and globally, yielding depth prediction at a higher level of detail and accuracy, allowing MonoViT to achieve state-of-the-art performance on the established KITTI dataset. Moreover, MonoViT proves its superior generalization capacities on other datasets such as Make3D and DrivingStereo.
Accepted by 3DV 2022
References in corpus (3)
Cited by in corpus (7)
- Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement
- MonoPP: Metric-Scaled Self-Supervised Monocular Depth Estimation by Planar-Parallax Geometry in Automotive Applications
- Digging into contrastive learning for robust depth estimation with diffusion models
- On Robust Cross-View Consistency in Self-Supervised Monocular Depth Estimation
- Improving Domain Generalization in Self-supervised Monocular Depth Estimation via Stabilized Adversarial Training
- TP3M: Transformer-based Pseudo 3D Image Matching with Reference Image
- MGNiceNet: Unified Monocular Geometric Scene Understanding