RegNet: Learning the Optimization of Direct Image-to-Image Pose Registration
arXiv:1812.10212
Abstract
Direct image-to-image alignment that relies on the optimization of photometric error metrics suffers from limited convergence range and sensitivity to lighting conditions. Deep learning approaches has been applied to address this problem by learning better feature representations using convolutional neural networks, yet still require a good initialization. In this paper, we demonstrate that the inaccurate numerical Jacobian limits the convergence range which could be improved greatly using learned approaches. Based on this observation, we propose a novel end-to-end network, RegNet, to learn the optimization of image-to-image pose registration. By jointly learning feature representation for each pixel and partial derivatives that replace handcrafted ones (e.g., numerical differentiation) in the optimization step, the neural network facilitates end-to-end optimization. The energy landscape is constrained on both the feature representation and the learned Jacobian, hence providing more flexibility for the optimization as a consequence leads to more robust and faster convergence. In a series of experiments, including a broad ablation study, we demonstrate that RegNet is able to converge for large-baseline image pairs with fewer iterations.
8 pages, 6 figures
References in corpus (10)
- Adam: A Method for Stochastic Optimization
- DeMoN: Depth and Motion Network for Learning Monocular Stereo
- Dilated Residual Networks
- Deep Image Homography Estimation
- BA-Net: Dense Bundle Adjustment Network
- LIFT: Learned Invariant Feature Transform
- CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
- LS-Net: Learning to Solve Nonlinear Least Squares for Monocular Stereo
- Relative Camera Pose Estimation Using Convolutional Neural Networks
- Aligning Across Large Gaps in Time