Unsupervised Learning of Depth and Ego-Motion from Video
arXiv:1704.07813
Abstract
We present an unsupervised learning framework for the task of monocular depth and camera motion estimation from unstructured video sequences. We achieve this by simultaneously training depth and camera pose estimation networks using the task of view synthesis as the supervisory signal. The networks are thus coupled via the view synthesis objective during training, but can be applied independently at test time. Empirical evaluation on the KITTI dataset demonstrates the effectiveness of our approach: 1) monocular depth performing comparably with supervised methods that use either ground-truth pose or depth for training, and 2) pose estimation performing favorably with established SLAM systems under comparable input settings.
Accepted to CVPR 2017. Project webpage: https://people.eecs.berkeley.edu/~tinghuiz/projects/SfMLearner/
References in corpus (7)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Deep Convolutional Inverse Graphics Network
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- Semi-Supervised Deep Learning for Monocular Depth Map Prediction
- DeepStereo: Learning to Predict New Views from the World's Imagery
- 3D Shape Induction from 2D Views of Multiple Objects
- gvnn: Neural Network Library for Geometric Computer Vision
Cited by in corpus (94)
- SfM-Net: Learning of Structure and Motion from Video
- Learning a Multi-View Stereo Machine
- Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination
- Self-supervised Learning of Motion Capture
- GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
- Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency
- Sparse-to-Dense: Depth Prediction from Sparse Depth Samples and a Single Image
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Unifying Map and Landmark Based Representations for Visual Navigation
- MegaDepth: Learning Single-View Depth Prediction from Internet Photos
- Unsupervised Video Object Segmentation for Deep Reinforcement Learning
- UnDeepVO: Monocular Visual Odometry through Unsupervised Deep Learning
- Toward Geometric Deep SLAM
- Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
- ST-GAN: Spatial Transformer Generative Adversarial Networks for Image Compositing
- Learning Depth from Monocular Videos using Direct Methods
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
- FutureMapping: The Computational Structure of Spatial AI Systems
- Unsupervised Learning of Dense Optical Flow, Depth and Egomotion from Sparse Event Data
- Self-Improving Visual Odometry
- Monocular Depth Estimation using Multi-Scale Continuous CRFs as Sequential Deep Networks
- DeepLO: Geometry-Aware Deep LiDAR Odometry
- Camera-based vehicle velocity estimation from monocular video
- HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation
- Transitive Invariance for Self-supervised Visual Representation Learning
- Unsupervised Depth Estimation, 3D Face Rotation and Replacement
- LEGO: Learning Edge with Geometry all at Once by Watching Videos
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- 3D Photography using Context-aware Layered Depth Inpainting
- Deep Stereo using Adaptive Thin Volume Representation with Uncertainty Awareness
- Semantic Flow for Fast and Accurate Scene Parsing
- Single View Stereo Matching
- RGBD-GAN: Unsupervised 3D Representation Learning From Natural Image Datasets via RGBD Image Synthesis
- DOVE: Learning Deformable 3D Objects by Watching Videos
- Joint Unsupervised Learning of Optical Flow and Depth by Watching Stereo Videos
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- Sequential Adversarial Learning for Self-Supervised Deep Visual Odometry
- Unsupervised learning of object frames by dense equivariant image labelling
- Disentangling Factors of Variation by Mixing Them
- SharinGAN: Combining Synthetic and Real Data for Unsupervised Geometry Estimation
- Unsupervised High-Resolution Depth Learning From Videos With Dual Networks
- X-Distill: Improving Self-Supervised Monocular Depth via Cross-Task Distillation
- Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments
- End-to-end depth from motion with stabilized monocular videos
- InverseRenderNet: Learning single image inverse rendering
- Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations
- Monocular Depth Estimators: Vulnerabilities and Attacks
- Do 2D GANs Know 3D Shape? Unsupervised 3D shape reconstruction from 2D Image GANs
- Visuomotor Understanding for Representation Learning of Driving Scenes
- PRGFlow: Benchmarking SWAP-Aware Unified Deep Visual Inertial Odometry
- Future Person Localization in First-Person Videos
- Layer-structured 3D Scene Inference via View Synthesis
- Unsupervised 3D Shape Learning from Image Collections in the Wild
- Semi-supervised Learning: Fusion of Self-supervised, Supervised Learning, and Multimodal Cues for Tactical Driver Behavior Detection
- Unsupervised Joint Learning of Depth, Optical Flow, Ego-motion from Video
- Structure-Attentioned Memory Network for Monocular Depth Estimation
- Unsupervised Video Depth Estimation Based on Ego-motion and Disparity Consensus
- Future Segmentation Using 3D Structure
- UnDEMoN 2.0: Improved Depth and Ego Motion Estimation through Deep Image Sampling
- Semantic Photometric Bundle Adjustment on Natural Sequences
- Unsupervised Event-based Learning of Optical Flow, Depth, and Egomotion
- S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation
- Unsupervised Representation Learning by Predicting Random Distances
- Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching
- Learning Multiple Dense Prediction Tasks from Partially Annotated Data
- Unsupervised Metric Relocalization Using Transform Consistency Loss
- Self-supervised Visual-LiDAR Odometry with Flip Consistency
- Generalizing to the Open World: Deep Visual Odometry with Online Adaptation
- Distill Knowledge from NRSfM for Weakly Supervised 3D Pose Learning
- SIGNet: Semantic Instance Aided Unsupervised 3D Geometry Perception
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- Learning by Inertia: Self-supervised Monocular Visual Odometry for Road Vehicles
- Multimodal Scale Consistency and Awareness for Monocular Self-Supervised Depth Estimation
- Adversarial View-Consistent Learning for Monocular Depth Estimation
- SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
- Why Convolutional Networks Learn Oriented Bandpass Filters: Theory and Empirical Support
- End-to-end Generative Zero-shot Learning via Few-shot Learning
- Robot Perception enables Complex Navigation Behavior via Self-Supervised Learning
- Region Deformer Networks for Unsupervised Depth Estimation from Unconstrained Monocular Videos
- Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space
- Multi-Hypothesis Visual-Inertial Flow
- Unsupervised monocular stereo matching
- Sequential Learning of Visual Tracking and Mapping Using Unsupervised Deep Neural Networks
- Improving Self-Supervised Single View Depth Estimation by Masking Occlusion
- Monocular Depth Estimation with Directional Consistency by Deep Networks
- Learning to Align Images using Weak Geometric Supervision
- A Multi-Task Learning Approach for Meal Assessment
- Learning Sports Camera Selection from Internet Videos
- Multi range Real-time depth inference from a monocular stabilized footage using a Fully Convolutional Neural Network
- Functionally Modular and Interpretable Temporal Filtering for Robust Segmentation
- Indoor GeoNet: Weakly Supervised Hybrid Learning for Depth and Pose Estimation
- Deep Weakly Supervised Positioning
- Does it work outside this benchmark? Introducing the Rigid Depth Constructor tool, depth validation dataset construction in rigid scenes for the masses
- LO-Net: Deep Real-time Lidar Odometry