From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
arXiv:1907.10326
Abstract
Estimating accurate depth from a single image is challenging because it is an ill-posed problem as infinitely many 3D scenes can be projected to the same 2D scene. However, recent works based on deep convolutional neural networks show great progress with plausible results. The convolutional neural networks are generally composed of two parts: an encoder for dense feature extraction and a decoder for predicting the desired depth. In the encoder-decoder schemes, repeated strided convolution and spatial pooling layers lower the spatial resolution of transitional outputs, and several techniques such as skip connections or multi-layer deconvolutional networks are adopted to recover the original resolution for effective dense prediction. In this paper, for more effective guidance of densely encoded features to the desired depth prediction, we propose a network architecture that utilizes novel local planar guidance layers located at multiple stages in the decoding phase. We show that the proposed method outperforms the state-of-the-art works with significant margin evaluating on challenging benchmarks. We also provide results from an ablation study to validate the effectiveness of the proposed method.
References in corpus (4)
Cited by in corpus (69)
- AdaBins: Depth Estimation using Adaptive Bins
- Cross-Modality Knowledge Distillation Network for Monocular 3D Object Detection
- Semantic Histogram Based Graph Matching for Real-Time Multi-Robot Global Localization in Large Scale Environment
- Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
- Transformers in Self-Supervised Monocular Depth Estimation with Unknown Camera Intrinsics
- On Deep Learning Techniques to Boost Monocular Depth Estimation for Autonomous Navigation
- Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement
- RefinedMPL: Refined Monocular PseudoLiDAR for 3D Object Detection in Autonomous Driving
- Adversarial Patch Attacks on Monocular Depth Estimation Networks
- Visual Attention-based Self-supervised Absolute Depth Estimation using Geometric Priors in Autonomous Driving
- Sparse-to-Continuous: Enhancing Monocular Depth Estimation using Occupancy Maps
- Less is More: Consistent Video Depth Estimation with Masked Frames Modeling
- Monocular Depth Estimation with Self-supervised Instance Adaptation
- SGTBN: Generating Dense Depth Maps from Single-Line LiDAR
- CutDepth:Edge-aware Data Augmentation in Depth Estimation
- Feature-metric Loss for Self-supervised Learning of Depth and Egomotion
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- Pyramid Frequency Network with Spatial Attention Residual Refinement Module for Monocular Depth Estimation
- Pyramid Feature Attention Network for Monocular Depth Prediction
- Fast and Accurate Single-Image Depth Estimation on Mobile Devices, Mobile AI 2021 Challenge: Report
- DwinFormer: Dual Window Transformers for End-to-End Monocular Depth Estimation
- OCM3D: Object-Centric Monocular 3D Object Detection
- NVS-MonoDepth: Improving Monocular Depth Prediction with Novel View Synthesis
- The Edge of Depth: Explicit Constraints between Segmentation and Depth
- ROBUSfT: Robust Real-Time Shape-from-Template, a C++ Library
- Bidirectional Attention Network for Monocular Depth Estimation
- Diffusion-Augmented Depth Prediction with Sparse Annotations
- Monocular Depth Estimators: Vulnerabilities and Attacks
- ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
- Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline
- Lidar Point Cloud Guided Monocular 3D Object Detection
- Ground-aware Monocular 3D Object Detection for Autonomous Driving
- Monocular Per-Object Distance Estimation with Masked Object Modeling
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- 3D-to-2D Distillation for Indoor Scene Parsing
- DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning
- Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration
- SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark under Multiple Environments
- Toward Hierarchical Self-Supervised Monocular Absolute Depth Estimation for Autonomous Driving Applications
- MVS2D: Efficient Multi-view Stereo via Attention-Driven 2D Convolutions
- SLURP: Side Learning Uncertainty for Regression Problems
- Edge-Guided Occlusion Fading Reduction for a Light-Weighted Self-Supervised Monocular Depth Estimation
- False Negative Reduction in Video Instance Segmentation using Uncertainty Estimates
- Calibrating Self-supervised Monocular Depth Estimation
- Robust Multi-Robot Global Localization with Unknown Initial Pose based on Neighbor Constraints
- Monocular Depth Decomposition of Semi-Transparent Volume Renderings
- Enhancing Monocular Depth Estimation with Multi-Source Auxiliary Tasks
- Learning Indoor Inverse Rendering with 3D Spatially-Varying Lighting
- Deep Multi Depth Panoramas for View Synthesis
- Monocular Depth Prediction through Continuous 3D Loss
- Self-Guided Instance-Aware Network for Depth Completion and Enhancement
- Radar-Camera Pixel Depth Association for Depth Completion
- Mirror3D: Depth Refinement for Mirror Surfaces
- Automatic Map Update Using Dashcam Videos
- Depth Estimation from Monocular Images and Sparse radar using Deep Ordinal Regression Network
- Does it work outside this benchmark? Introducing the Rigid Depth Constructor tool, depth validation dataset construction in rigid scenes for the masses
- How Much Depth Information can Radar Contribute to a Depth Estimation Model?
- SM3D: Simultaneous Monocular Mapping and 3D Detection
- MonoPLFlowNet: Permutohedral Lattice FlowNet for Real-Scale 3D Scene FlowEstimation with Monocular Images
- Towards Comprehensive Monocular Depth Estimation: Multiple Heads Are Better Than One
- Topological Regularization for Dense Prediction
- Facial Depth and Normal Estimation using Single Dual-Pixel Camera
- Geometry Enhancements from Visual Content: Going Beyond Ground Truth
- Predicting Depth from Semantic Segmentation using Game Engine Dataset
- Learning Monocular 3D Vehicle Detection without 3D Bounding Box Labels
- Weakly-Supervised Monocular Depth Estimationwith Resolution-Mismatched Data
- -Cal: Calibrated aleatoric uncertainty estimation from neural networks for robot perception
- Polarimetric Monocular Dense Mapping Using Relative Deep Depth Prior
- Improved Point Transformation Methods For Self-Supervised Depth Prediction