Multi-Scale Continuous CRFs as Sequential Deep Networks for Monocular Depth Estimation
arXiv:1704.02157
Abstract
This paper addresses the problem of depth estimation from a single still image. Inspired by recent works on multi- scale convolutional neural networks (CNN), we propose a deep model which fuses complementary information derived from multiple CNN side outputs. Different from previous methods, the integration is obtained by means of continuous Conditional Random Fields (CRFs). In particular, we propose two different variations, one based on a cascade of multiple CRFs, the other on a unified graphical model. By designing a novel CNN implementation of mean-field updates for continuous CRFs, we show that both proposed models can be regarded as sequential deep networks and that training can be performed end-to-end. Through extensive experimental evaluation we demonstrate the effective- ness of the proposed approach and establish new state of the art results on publicly available datasets.
Accepted as a spotlight paper at CVPR 2017
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Learning Deconvolution Network for Semantic Segmentation
- DepthTransfer: Depth Extraction from Video Using Non-parametric Sampling
- Learning Cross-Modal Deep Representations for Robust Pedestrian Detection
Cited by in corpus (15)
- Peeking Behind Objects: Layered Depth Prediction from a Single Image
- Exploiting the Potential of Standard Convolutional Autoencoders for Image Restoration by Evolutionary Search
- Monocular Depth Estimation using Multi-Scale Continuous CRFs as Sequential Deep Networks
- Camera-based vehicle velocity estimation from monocular video
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- Deep attention-based classification network for robust depth prediction
- Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments
- Unsupervised High-Resolution Depth Learning From Videos With Dual Networks
- Monocular Depth Estimation with Hierarchical Fusion of Dilated CNNs and Soft-Weighted-Sum Inference
- SymmNet: A Symmetric Convolutional Neural Network for Occlusion Detection
- S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation
- Attentional Separation-and-Aggregation Network for Self-supervised Depth-Pose Learning in Dynamic Scenes
- CI-Net: Contextual Information for Joint Semantic Segmentation and Depth Estimation
- Self-Supervised Monocular Image Depth Learning and Confidence Estimation
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views