Unsupervised learning of depth and motion
arXiv:1312.3429
Abstract
We present a model for the joint estimation of disparity and motion. The model is based on learning about the interrelations between images from multiple cameras, multiple frames in a video, or the combination of both. We show that learning depth and motion cues, as well as their combinations, from data is possible within a single type of architecture and a single type of learning algorithm, by using biologically inspired "complex cell" like units, which encode correlations between the pixels across image pairs. Our experimental results show that the learning of depth and motion makes it possible to achieve state-of-the-art performance in 3-D activity analysis, and to outperform existing hand-engineered 3-D motion features by a very large margin.
Cited by in corpus (8)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- FlowNet: Learning Optical Flow with Convolutional Networks
- DeepStereo: Learning to Predict New Views from the World's Imagery
- Learning to Extract Motion from Videos in Convolutional Neural Networks
- Evolution of Visual Odometry Techniques
- End-to-end depth from motion with stabilized monocular videos
- Neural Network Regularization via Robust Weight Factorization
- AD-VO: Scale-Resilient Visual Odometry Using Attentive Disparity Map