Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields
arXiv:1502.07411 · doi:10.1109/TPAMI.2015.2505283
Abstract
In this article, we tackle the problem of depth estimation from single monocular images. Compared with depth estimation using multiple images such as stereo depth perception, depth from monocular images is much more challenging. Prior work typically focuses on exploiting geometric priors or additional sources of information, most using hand-crafted features. Recently, there is mounting evidence that features from deep convolutional neural networks (CNN) set new records for various vision applications. On the other hand, considering the continuous characteristic of the depth values, depth estimations can be naturally formulated as a continuous conditional random field (CRF) learning problem. Therefore, here we present a deep convolutional neural field model for estimating depths from single monocular images, aiming to jointly explore the capacity of deep CNN and continuous CRF. In particular, we propose a deep structured learning scheme which learns the unary and pairwise potentials of continuous CRF in a unified deep CNN framework. We then further propose an equally effective model based on fully convolutional networks and a novel superpixel pooling method, which is times faster, to speedup the patch-wise convolutions in the deep model. With this more efficient model, we are able to design deeper networks to pursue better performance. Experiments on both indoor and outdoor scene datasets demonstrate that the proposed method outperforms state-of-the-art depth estimation approaches.
Appearing in IEEE T. Pattern Analysis and Machine Intelligence. Journal version of arXiv:1411.6387 . Test code is available at https://bitbucket.org/fayao/dcnf-fcsp
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fully Convolutional Networks for Semantic Segmentation
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation
- DepthTransfer: Depth Extraction from Video Using Non-parametric Sampling
- Combining the Best of Graphical Models and ConvNets for Semantic Segmentation
Cited by in corpus (179)
- FCOS: Fully Convolutional One-Stage Object Detection
- DeMoN: Depth and Motion Network for Learning Monocular Stereo
- AdaBins: Depth Estimation using Adaptive Bins
- From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
- A Survey on Deep Learning Techniques for Stereo-based Depth Estimation
- Monocular Depth Estimation Based On Deep Learning: An Overview
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer
- Unsupervised Reverse Domain Adaptation for Synthetic Medical Images via Adversarial Training
- Unsupervised Scale-consistent Depth Learning from Video
- Synthetic Depth-of-Field with a Single-Camera Mobile Phone
- Adaptive Context-Aware Multi-Modal Network for Depth Completion
- Perception and Navigation in Autonomous Systems in the Era of Learning: A Survey
- GridDehazeNet: Attention-Based Multi-Scale Network for Image Dehazing
- GANVO: Unsupervised Deep Monocular Visual Odometry and Depth Estimation with Generative Adversarial Networks
- Soft Rasterizer: A Differentiable Renderer for Image-based 3D Reasoning
- Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
- UniFuse: Unidirectional Fusion for 360 Panorama Depth Estimation
- Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
- A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence
- Approaches, Challenges, and Applications for Deep Visual Odometry: Toward to Complicated and Emerging Areas
- Detail Preserving Depth Estimation from a Single Image Using Attention Guided Networks
- Using User Generated Online Photos to Estimate and Monitor Air Pollution in Major Cities
- Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction
- Two-shot Spatially-varying BRDF and Shape Estimation
- GeoNet++: Iterative Geometric Neural Network with Edge-Aware Refinement for Joint Depth and Surface Normal Estimation
- Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
- Unsupervised Learning of Depth, Optical Flow and Pose with Occlusion from 3D Geometry
- Analyzing Modular CNN Architectures for Joint Depth Prediction and Semantic Segmentation
- Multi-Scale Feature Fusion: Learning Better Semantic Segmentation for Road Pothole Detection
- On Deep Learning Techniques to Boost Monocular Depth Estimation for Autonomous Navigation
- J-MOD: Joint Monocular Obstacle Detection and Depth Estimation
- Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement
- Unsupervised Monocular Depth Learning in Dynamic Scenes
- Enforcing geometric constraints of virtual normal for depth prediction
- Driving Scene Perception Network: Real-time Joint Detection, Depth Estimation and Semantic Segmentation
- MiniNet: An extremely lightweight convolutional neural network for real-time unsupervised monocular depth estimation
- DIODE: A Dense Indoor and Outdoor DEpth Dataset
- Deep Learning Based 3D Segmentation: A Survey
- Adversarial Patch Attacks on Monocular Depth Estimation Networks
- Deep Learning with Cinematic Rendering: Fine-Tuning Deep Neural Networks Using Photorealistic Medical Images
- Rethinking Monocular Depth Estimation with Adversarial Training
- FCOS: A simple and strong anchor-free object detector
- 3G structure for image caption generation
- RSGM: Real-time Raster-Respecting Semi-Global Matching for Power-Constrained Systems
- Just-in-Time Reconstruction: Inpainting Sparse Maps using Single View Depth Predictors as Priors
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- Benchmarking Single Image Dehazing and Beyond
- Joint Self-supervised Depth and Optical Flow Estimation towards Dynamic Objects
- Deeply Learning the Messages in Message Passing Inference
- Learning Robotic Navigation from Experience: Principles, Methods, and Recent Results
- Motion-based Camera Localization System in Colonoscopy Videos
- A feature-supervised generative adversarial network for environmental monitoring during hazy days
- Unsupervised Learning of Monocular Depth and Ego-Motion Using Multiple Masks
- DF-VO: What Should Be Learnt for Visual Odometry?
- Conditional Convolutions for Instance Segmentation
- A Survey on Deep Learning Architectures for Image-based Depth Reconstruction
- Como funciona o Deep Learning
- SimCol3D -- 3D Reconstruction during Colonoscopy Challenge
- Discriminative Training of Deep Fully-connected Continuous CRF with Task-specific Loss
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- Self-Supervised Joint Learning Framework of Depth Estimation via Implicit Cues
- Depth from Monocular Images using a Semi-Parallel Deep Neural Network (SPDNN) Hybrid Architecture
- Pattern-Affinitive Propagation across Depth, Surface Normal and Semantic Segmentation
- Multi-Camera Collaborative Depth Prediction via Consistent Structure Estimation
- Lightweight Monocular Depth Estimation Model by Joint End-to-End Filter pruning
- Self-supervised learning for autonomous vehicles perception: A conciliation between analytical and learning methods
- On Incremental Structure-from-Motion using Lines
- DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene from Sparse LiDAR Data and Single Color Image
- Structured Knowledge Distillation for Dense Prediction
- Pyramid Frequency Network with Spatial Attention Residual Refinement Module for Monocular Depth Estimation
- Uncovering local aggregated air quality index with smartphone captured images leveraging efficient deep convolutional neural network
- Unsupervised Learning of Depth and Ego-Motion from Cylindrical Panoramic Video
- Adversarial Attacks on Monocular Depth Estimation
- Unsupervised Neural Rendering for Image Hazing
- Fast and Accurate Single-Image Depth Estimation on Mobile Devices, Mobile AI 2021 Challenge: Report
- Sequential Adversarial Learning for Self-Supervised Deep Visual Odometry
- Exploiting temporal consistency for real-time video depth estimation
- Learning monocular depth estimation infusing traditional stereo knowledge
- Online Adaptation through Meta-Learning for Stereo Depth Estimation
- Deep Learning-Based 3D Instance and Semantic Segmentation: A Review
- PNet: Patch-match and Plane-regularization for Unsupervised Indoor Depth Estimation
- Double Refinement Network for Efficient Indoor Monocular Depth Estimation
- SharinGAN: Combining Synthetic and Real Data for Unsupervised Geometry Estimation
- Visual Odometry Revisited: What Should Be Learnt?
- Geometry-Aware Symmetric Domain Adaptation for Monocular Depth Estimation
- Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments
- Consistent Video Depth Estimation
- Joint Prediction of Monocular Depth and Structure using Planar and Parallax Geometry
- Self-Supervised Learning of Depth and Ego-motion with Differentiable Bundle Adjustment
- Bidirectional Attention Network for Monocular Depth Estimation
- Task-Aware Monocular Depth Estimation for 3D Object Detection
- Bi-Real Net: Binarizing Deep Network Towards Real-Network Performance
- Self-supervised Learning for Single View Depth and Surface Normal Estimation
- MonoPP: Metric-Scaled Self-Supervised Monocular Depth Estimation by Planar-Parallax Geometry in Automotive Applications
- LiDAR Data Enrichment Using Deep Learning Based on High-Resolution Image: An Approach to Achieve High-Performance LiDAR SLAM Using Low-cost LiDAR
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- MonSter: Awakening the Mono in Stereo
- On the Importance of Stereo for Accurate Depth Estimation: An Efficient Semi-Supervised Deep Neural Network Approach
- DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular Videos
- Monocular Per-Object Distance Estimation with Masked Object Modeling
- PlaneRCNN: 3D Plane Detection and Reconstruction from a Single Image
- Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation
- Deep Shape-from-Template: Wide-Baseline, Dense and Fast Registration and Deformable Reconstruction from a Single Image
- Auxiliary Learning for Deep Multi-task Learning
- Realistic Large-Scale Fine-Depth Dehazing Dataset from 3D Videos
- Unsupervised Learning of Depth and Ego-Motion from Cylindrical Panoramic Video with Applications for Virtual Reality
- Pseudo Supervised Monocular Depth Estimation with Teacher-Student Network
- Recognizing Image Objects by Relational Analysis Using Heterogeneous Superpixels and Deep Convolutional Features
- Deep Learning in Robotics: A Review of Recent Research
- Monocular Depth Estimation with Augmented Ordinal Depth Relationships
- End-to-End Learnable Geometric Vision by Backpropagating PnP Optimization
- EndoSLAM Dataset and An Unsupervised Monocular Visual Odometry and Depth Estimation Approach for Endoscopic Videos: Endo-SfMLearner
- Learning Depth from Monocular Videos Using Synthetic Data: A Temporally-Consistent Domain Adaptation Approach
- A Framework for 3D Tracking of Frontal Dynamic Objects in Autonomous Cars
- Learning to Recover 3D Scene Shape from a Single Image
- Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration
- MobileDepth: Efficient Monocular Depth Prediction on Mobile Devices
- Towards Domain-agnostic Depth Completion
- SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark under Multiple Environments
- Toward Hierarchical Self-Supervised Monocular Absolute Depth Estimation for Autonomous Driving Applications
- DeepSFM: Structure From Motion Via Deep Bundle Adjustment
- Structure-Attentioned Memory Network for Monocular Depth Estimation
- Learning Object-specific Distance from a Monocular Image
- Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image
- Geo-Supervised Visual Depth Prediction
- Edge-Guided Occlusion Fading Reduction for a Light-Weighted Self-Supervised Monocular Depth Estimation
- Semi-supervised Learning for Few-shot Image-to-Image Translation
- Calibrating Self-supervised Monocular Depth Estimation
- FIS-Nets: Full-image Supervised Networks for Monocular Depth Estimation
- Virtual Normal: Enforcing Geometric Constraints for Accurate and Robust Depth Prediction
- Generating and Exploiting Probabilistic Monocular Depth Estimates
- S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation
- Instance-wise Depth and Motion Learning from Monocular Videos
- Learning Joint 2D-3D Representations for Depth Completion
- Deep Animation Video Interpolation in the Wild
- Structured Depth Prediction in Challenging Monocular Video Sequences
- Enhancing Monocular Height Estimation via Sparse LiDAR-Guided Correction
- ODE-CNN: Omnidirectional Depth Extension Networks
- FusionMapping: Learning Depth Prediction with Monocular Images and 2D Laser Scans
- Inferring Distributions Over Depth from a Single Image
- Robust Consistent Video Depth Estimation
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- Unsupervised Learning of Camera Pose with Compositional Re-estimation
- Mix and match networks: cross-modal alignment for zero-pair image-to-image translation
- Self-Supervised Learning of Depth and Motion Under Photometric Inconsistency
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- Targeted Adversarial Perturbations for Monocular Depth Prediction
- Increased-Range Unsupervised Monocular Depth Estimation
- Net: Semantic-Aware Self-supervised Depth Estimation with Monocular Videos and Synthetic Data
- Deep Classification Network for Monocular Depth Estimation
- StructDepth: Leveraging the structural regularities for self-supervised indoor depth estimation
- Unsupervised Monocular Depth Perception: Focusing on Moving Objects
- Camera Pose Matters: Improving Depth Prediction by Mitigating Pose Distribution Bias
- Defocus Map Estimation and Deblurring from a Single Dual-Pixel Image
- Visual Relationship Prediction via Label Clustering and Incorporation of Depth Information
- DAN: A Deformation-Aware Network for Consecutive Biomedical Image Interpolation
- Sparse2Dense: From direct sparse odometry to dense 3D reconstruction
- On the Synergies between Machine Learning and Binocular Stereo for Depth Estimation from Images: a Survey
- Crowdsourced 3D Mapping: A Combined Multi-View Geometry and Self-Supervised Learning Approach
- n-MeRCI: A new Metric to Evaluate the Correlation Between Predictive Uncertainty and True Error
- InsertionNet -- A Scalable Solution for Insertion
- Towards Comprehensive Monocular Depth Estimation: Multiple Heads Are Better Than One
- MonoPLFlowNet: Permutohedral Lattice FlowNet for Real-Scale 3D Scene FlowEstimation with Monocular Images
- Late or Earlier Information Fusion from Depth and Spectral Data? Large-Scale Digital Surface Model Refinement by Hybrid-cGAN
- Synthesizing Photorealistic Images with Deep Generative Learning
- Unsupervised Learning of Depth and Depth-of-Field Effect from Natural Images with Aperture Rendering Generative Adversarial Networks
- Self-Supervised Learning of Depth and Ego-Motion from Video by Alternative Training and Geometric Constraints from 3D to 2D
- Weakly-Supervised Monocular Depth Estimationwith Resolution-Mismatched Data
- Relational Neural Markov Random Fields
- Learning to Predict the 3D Layout of a Scene
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views
- End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular Camera
- Adversarial Structure Matching for Structured Prediction Tasks
- CrDoCo: Pixel-level Domain Transfer with Cross-Domain Consistency
- Unsupervised monocular stereo matching
- Unstructured Multi-View Depth Estimation Using Mask-Based Multiplane Representation
- Veritatem Dies Aperit- Temporally Consistent Depth Prediction Enabled by a Multi-Task Geometric and Semantic Scene Understanding Approach
- When Autonomous Systems Meet Accuracy and Transferability through AI: A Survey