VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization
arXiv:1702.06521
Abstract
Machine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of the proposed learning-based approaches exploit the valuable constraint of temporal smoothness, often leading to situations where the per-frame error is larger than the camera motion. In this paper we propose a recurrent model for performing 6-DoF localization of video-clips. We find that, even by considering only short sequences (20 frames), the pose estimates are smoothed and the localization error can be drastically reduced. Finally, we consider means of obtaining probabilistic pose estimates from our model. We evaluate our method on openly-available real-world autonomous driving and indoor localization datasets.
To appear at CVPR 2017
References in corpus (2)
Cited by in corpus (10)
- Towards Monocular Vision based Obstacle Avoidance through Deep Reinforcement Learning
- Active Neural Localization
- Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions
- LS-Net: Learning to Solve Nonlinear Least Squares for Monocular Stereo
- Learning models for visual 3D localization with implicit mapping
- To Learn or Not to Learn: Visual Localization from Essential Matrices
- KFNet: Learning Temporal Camera Relocalization using Kalman Filtering
- PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning
- Improving drone localisation around wind turbines using monocular model-based tracking
- Deep Weakly Supervised Positioning