DeepV2D: Video to Depth with Differentiable Structure from Motion
arXiv:1812.04605
Abstract
We propose DeepV2D, an end-to-end deep learning architecture for predicting depth from video. DeepV2D combines the representation ability of neural networks with the geometric principles governing image formation. We compose a collection of classical geometric algorithms, which are converted into trainable modules and combined into an end-to-end differentiable architecture. DeepV2D interleaves two stages: motion estimation and depth estimation. During inference, motion and depth estimation are alternated and converge to accurate depth. Code is available https://github.com/princeton-vl/DeepV2D.
References in corpus (9)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- DeepVO: Towards End-to-End Visual Odometry with Deep Recurrent Convolutional Neural Networks
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- SfM-Net: Learning of Structure and Motion from Video
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- High Quality Monocular Depth Estimation via Transfer Learning
- BA-Net: Dense Bundle Adjustment Network
- Semi-Supervised Deep Learning for Monocular Depth Map Prediction
- LS-Net: Learning to Solve Nonlinear Least Squares for Monocular Stereo
Cited by in corpus (18)
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- On Deep Learning Techniques to Boost Monocular Depth Estimation for Autonomous Navigation
- A Survey of Simultaneous Localization and Mapping with an Envision in 6G Wireless Networks
- DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost Volume
- Less is More: Consistent Video Depth Estimation with Masked Frames Modeling
- Space-time Neural Irradiance Fields for Free-Viewpoint Video
- Two Stream Networks for Self-Supervised Ego-Motion Estimation
- PNet: Patch-match and Plane-regularization for Unsupervised Indoor Depth Estimation
- NVS-MonoDepth: Improving Monocular Depth Prediction with Novel View Synthesis
- Consistent Video Depth Estimation
- Self-Supervised 3D Keypoint Learning for Ego-motion Estimation
- FADEC: FPGA-based Acceleration of Video Depth Estimation by HW/SW Co-design
- Monocular 3D Object Detection: An Extrinsic Parameter Free Approach
- DeepSFM: Structure From Motion Via Deep Bundle Adjustment
- Robust Consistent Video Depth Estimation
- Tangent Space Backpropagation for 3D Transformation Groups
- Video Depth Estimation by Fusing Flow-to-Depth Proposals