Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
arXiv:1603.04992
Abstract
A significant weakness of most current deep Convolutional Neural Networks is the need to train them using vast amounts of manu- ally labelled data. In this work we propose a unsupervised framework to learn a deep convolutional neural network for single view depth predic- tion, without requiring a pre-training stage or annotated ground truth depths. We achieve this by training the network in a manner analogous to an autoencoder. At training time we consider a pair of images, source and target, with small, known camera motion between the two such as a stereo pair. We train the convolutional encoder for the task of predicting the depth map for the source image. To do so, we explicitly generate an inverse warp of the target image using the predicted depth and known inter-view displacement, to reconstruct the source image; the photomet- ric error in the reconstruction is the reconstruction loss for the encoder. The acquisition of this training data is considerably simpler than for equivalent systems, requiring no manual annotation, nor calibration of depth sensor to camera. We show that our network trained on less than half of the KITTI dataset (without any further augmentation) gives com- parable performance to that of the state of art supervised methods for single view depth estimation.
Accepted for publication at ECCV 2016
References in corpus (7)
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Learning Deconvolution Network for Semantic Segmentation
- SceneNet: Understanding Real World Indoor Scenes With Synthetic Data
- Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation
- Deep3D: Fully Automatic 2D-to-3D Video Conversion with Deep Convolutional Neural Networks
Cited by in corpus (153)
- A Deep Learning Framework for Unsupervised Affine and Deformable Image Registration
- From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
- SfM-Net: Learning of Structure and Motion from Video
- Monocular Depth Estimation Based On Deep Learning: An Overview
- Unsupervised Learning of Depth and Ego-Motion from Video
- Learning a Multi-View Stereo Machine
- Unsupervised Scale-consistent Depth Learning from Video
- Self-Supervised Learning for Stereo Matching with Self-Improving Ability
- Self-supervised Learning of Motion Capture
- Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
- GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
- Single Image Depth Estimation: An Overview
- A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence
- Sparse-to-Dense: Depth Prediction from Sparse Depth Samples and a Single Image
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Masked GANs for Unsupervised Depth and Pose Prediction with Scale Consistency
- Learning Unsupervised Multi-View Stereopsis via Robust Photometric Consistency
- Unsupervised Learning of Depth, Optical Flow and Pose with Occlusion from 3D Geometry
- Privacy Intelligence: A Survey on Image Privacy in Online Social Networks
- Peeking Behind Objects: Layered Depth Prediction from a Single Image
- MegaDepth: Learning Single-View Depth Prediction from Internet Photos
- On Deep Learning Techniques to Boost Monocular Depth Estimation for Autonomous Navigation
- Unsupervised Monocular Depth Learning in Dynamic Scenes
- MiniNet: An extremely lightweight convolutional neural network for real-time unsupervised monocular depth estimation
- DIODE: A Dense Indoor and Outdoor DEpth Dataset
- UnDeepVO: Monocular Visual Odometry through Unsupervised Deep Learning
- Recurrent Neural Network for (Un-)supervised Learning of Monocular VideoVisual Odometry and Depth
- Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- DF-SLAM: A Deep-Learning Enhanced Visual SLAM System based on Deep Local Features
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- SAFENet: Self-Supervised Monocular Depth Estimation with Semantic-Aware Feature Extraction
- Unsupervised Learning of Dense Optical Flow, Depth and Egomotion from Sparse Event Data
- Monocular Depth Estimation using Multi-Scale Continuous CRFs as Sequential Deep Networks
- Data-Efficient Learning for Sim-to-Real Robotic Grasping using Deep Point Cloud Prediction Networks
- On the Benefit of Adversarial Training for Monocular Depth Estimation
- DF-VO: What Should Be Learnt for Visual Odometry?
- Motion-based Camera Localization System in Colonoscopy Videos
- Back to Basics: Unsupervised Learning of Optical Flow via Brightness Constancy and Motion Smoothness
- Camera-based vehicle velocity estimation from monocular video
- Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
- HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation
- Unsupervised Depth Estimation, 3D Face Rotation and Replacement
- Learning 3D Dynamic Scene Representations for Robot Manipulation
- ALIGNet: Partial-Shape Agnostic Alignment via Unsupervised Learning
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single-View Depth Estimation with Neural Positional Encoding and Distilled Matting Loss
- On Incremental Structure-from-Motion using Lines
- Deep Stereo using Adaptive Thin Volume Representation with Uncertainty Awareness
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Single View Stereo Matching
- Joint Unsupervised Learning of Optical Flow and Depth by Watching Stereo Videos
- Robustness Meets Deep Learning: An End-to-End Hybrid Pipeline for Unsupervised Learning of Egomotion
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- DOVE: Learning Deformable 3D Objects by Watching Videos
- Sequential Adversarial Learning for Self-Supervised Deep Visual Odometry
- Learning monocular depth estimation infusing traditional stereo knowledge
- Online Adaptation through Meta-Learning for Stereo Depth Estimation
- Unsupervised learning of object frames by dense equivariant image labelling
- Visual Odometry Revisited: What Should Be Learnt?
- Weakly- and Self-Supervised Learning for Content-Aware Deep Image Retargeting
- Inferring Point Clouds from Single Monocular Images by Depth Intermediation
- SharinGAN: Combining Synthetic and Real Data for Unsupervised Geometry Estimation
- FoveaNet: Perspective-aware Urban Scene Parsing
- Dense Depth Posterior (DDP) from Single Image and Sparse Range
- The Edge of Depth: Explicit Constraints between Segmentation and Depth
- Nonlinear Power Method for Computing Eigenvectors of Proximal Operators and Neural Networks
- Geometry-Aware Symmetric Domain Adaptation for Monocular Depth Estimation
- DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation
- EyeNet: A Multi-Task Network for Off-Axis Eye Gaze Estimation and User Understanding
- Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution
- Self-Supervised Learning of Depth and Ego-motion with Differentiable Bundle Adjustment
- ES-Net: An Efficient Stereo Matching Network
- End-to-end depth from motion with stabilized monocular videos
- Monocular Depth Estimation with Hierarchical Fusion of Dilated CNNs and Soft-Weighted-Sum Inference
- InverseRenderNet: Learning single image inverse rendering
- Towards Better Generalization: Joint Depth-Pose Learning without PoseNet
- Spherical View Synthesis for Self-Supervised 360 Depth Estimation
- Deep Learning based Monocular Depth Prediction: Datasets, Methods and Applications
- 3D Reconstruction of Simple Objects from A Single View Silhouette Image
- Learning Independent Object Motion from Unlabelled Stereoscopic Videos
- Self-supervised Learning for Single View Depth and Surface Normal Estimation
- DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular Videos
- Deep Depth From Focus
- Web Stereo Video Supervision for Depth Prediction from Dynamic Scenes
- Domain Adaptive Semantic Segmentation with Self-Supervised Depth Estimation
- On the Importance of Stereo for Accurate Depth Estimation: An Efficient Semi-Supervised Deep Neural Network Approach
- MonSter: Awakening the Mono in Stereo
- Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation
- Layer-structured 3D Scene Inference via View Synthesis
- Synthesizing a 4D Spatio-Angular Consistent Light Field from a Single Image
- A Novel Monocular Disparity Estimation Network with Domain Transformation and Ambiguity Learning
- Using Simulated Data to Generate Images of Climate Change
- Driven to Distraction: Self-Supervised Distractor Learning for Robust Monocular Visual Odometry in Urban Environments
- View Extrapolation of Human Body from a Single Image
- Monocular Depth Estimation with Augmented Ordinal Depth Relationships
- Learning Object-specific Distance from a Monocular Image
- Unsupervised Video Depth Estimation Based on Ego-motion and Disparity Consensus
- Structure-Attentioned Memory Network for Monocular Depth Estimation
- Geo-Supervised Visual Depth Prediction
- On Machine Learning and Structure for Mobile Robots
- RGB-D Local Implicit Function for Depth Completion of Transparent Objects
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- Synthetic Defocus and Look-Ahead Autofocus for Casual Videography
- DeepVIO: Self-supervised Deep Learning of Monocular Visual Inertial Odometry using 3D Geometric Constraints
- UnDEMoN 2.0: Improved Depth and Ego Motion Estimation through Deep Image Sampling
- Pano Popups: Indoor 3D Reconstruction with a Plane-Aware Network
- Virtual Normal: Enforcing Geometric Constraints for Accurate and Robust Depth Prediction
- Towards CNN Map Compression for camera relocalisation
- Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching
- Learning Joint 2D-3D Representations for Depth Completion
- TransCamP: Graph Transformer for 6-DoF Camera Pose Estimation
- S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation
- EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion Segmentation
- Calibrating Self-supervised Monocular Depth Estimation
- Instance-wise Depth and Motion Learning from Monocular Videos
- Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos
- Unsupervised Learning of Camera Pose with Compositional Re-estimation
- Learning what is above and what is below: horizon approach to monocular obstacle detection
- Self-Supervised Learning of Depth and Motion Under Photometric Inconsistency
- Depth Completion via Deep Basis Fitting
- Generalizing to the Open World: Deep Visual Odometry with Online Adaptation
- 5D Light Field Synthesis from a Monocular Video
- StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo Matching
- Self-Supervised Human Depth Estimation from Monocular Videos
- Deep feature fusion for self-supervised monocular depth prediction
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- LiStereo: Generate Dense Depth Maps from LIDAR and Stereo Imagery
- SIGNet: Semantic Instance Aided Unsupervised 3D Geometry Perception
- SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings
- Increased-Range Unsupervised Monocular Depth Estimation
- Region Deformer Networks for Unsupervised Depth Estimation from Unconstrained Monocular Videos
- Relative Depth Estimation as a Ranking Problem
- RRNet: Repetition-Reduction Network for Energy Efficient Decoder of Depth Estimation
- Learning Single-Image Depth from Videos using Quality Assessment Networks
- Multimodal Scale Consistency and Awareness for Monocular Self-Supervised Depth Estimation
- Unsupervised Monocular Depth Perception: Focusing on Moving Objects
- Monocular Depth Prediction through Continuous 3D Loss
- Is Depth Really Necessary for Salient Object Detection?
- Indoor GeoNet: Weakly Supervised Hybrid Learning for Depth and Pose Estimation
- Multi range Real-time depth inference from a monocular stabilized footage using a Fully Convolutional Neural Network
- Improving Self-Supervised Single View Depth Estimation by Masking Occlusion
- Depth Assisted Full Resolution Network for Single Image-based View Synthesis
- Learning Rank Reduced Interpolation with Principal Component Analysis
- Single Image Depth Prediction with Wavelet Decomposition
- Does it work outside this benchmark? Introducing the Rigid Depth Constructor tool, depth validation dataset construction in rigid scenes for the masses
- Multi-Hypothesis Visual-Inertial Flow
- Unsupervised monocular stereo matching
- Unsupervised Cross-spectral Stereo Matching by Learning to Synthesize
- Monocular Depth Estimation with Directional Consistency by Deep Networks
- Implicit Label Augmentation on Partially Annotated Clips via Temporally-Adaptive Features Learning