Deep Ordinal Regression Network for Monocular Depth Estimation
arXiv:1806.02446
Abstract
Monocular depth estimation, which plays a crucial role in understanding 3D scene geometry, is an ill-posed problem. Recent methods have gained significant improvement by exploring image-level information and hierarchical features from deep convolutional neural networks (DCNNs). These methods model depth estimation as a regression problem and train the regression networks by minimizing mean squared error, which suffers from slow convergence and unsatisfactory local solutions. Besides, existing depth estimation networks employ repeated spatial pooling operations, resulting in undesirable low-resolution feature maps. To obtain high-resolution depth maps, skip-connections or multi-layer deconvolution networks are required, which complicates network training and consumes much more computations. To eliminate or at least largely reduce these problems, we introduce a spacing-increasing discretization (SID) strategy to discretize depth and recast depth network learning as an ordinal regression problem. By training the network using an ordinary regression loss, our method achieves much higher accuracy and \dd{faster convergence in synch}. Furthermore, we adopt a multi-scale network structure which avoids unnecessary spatial pooling and captures multi-scale information in parallel. The method described in this paper achieves state-of-the-art results on four challenging benchmarks, i.e., KITTI [17], ScanNet [9], Make3D [50], and NYU Depth v2 [42], and win the 1st prize in Robust Vision Challenge 2018. Code has been made available at: https://github.com/hufu6371/DORN.
CVPR 2018
Cited by in corpus (30)
- The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation
- Monocular Depth Estimation: A Survey
- Progressive Coordinate Transforms for Monocular 3D Object Detection
- Monocular 3D Object Detection with Pseudo-LiDAR Point Cloud
- SAFENet: Self-Supervised Monocular Depth Estimation with Semantic-Aware Feature Extraction
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object Detection
- Monocular Depth Estimation with Self-supervised Instance Adaptation
- Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps with Accurate Object Boundaries
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- Deep attention-based classification network for robust depth prediction
- ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving
- Unsupervised High-Resolution Depth Learning From Videos With Dual Networks
- Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments
- InverseRenderNet: Learning single image inverse rendering
- Cascaded channel pruning using hierarchical self-distillation
- Ground-aware Monocular 3D Object Detection for Autonomous Driving
- Structure-Attentioned Memory Network for Monocular Depth Estimation
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- Calibrating Self-supervised Monocular Depth Estimation
- Attentional Separation-and-Aggregation Network for Self-supervised Depth-Pose Learning in Dynamic Scenes
- S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation
- SIGNet: Semantic Instance Aided Unsupervised 3D Geometry Perception
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- DAN: A Deformation-Aware Network for Consecutive Biomedical Image Interpolation
- GrabAR: Occlusion-aware Grabbing Virtual Objects in AR
- SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
- Learning Single-Image Depth from Videos using Quality Assessment Networks
- Polarimetric Monocular Dense Mapping Using Relative Deep Depth Prior
- Balanced Depth Completion between Dense Depth Inference and Sparse Range Measurements via KISS-GP
- Single Image Depth Prediction with Wavelet Decomposition