LIFT: Learned Invariant Feature Transform
arXiv:1603.09114
Abstract
We introduce a novel Deep Network architecture that implements the full feature point handling pipeline, that is, detection, orientation estimation, and feature description. While previous works have successfully tackled each one of these problems individually, we show how to learn to do all three in a unified manner while preserving end-to-end differentiability. We then demonstrate that our Deep pipeline outperforms state-of-the-art methods on a number of benchmark datasets, without the need of retraining.
Accepted to ECCV 2016 (spotlight)
References in corpus (3)
Cited by in corpus (39)
- 3DFeat-Net: Weakly Supervised Local 3D Features for Point Cloud Registration
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Universal Correspondence Network
- Human Pose Regression by Combining Indirect Part Detection and Contextual Information
- Numerical Coordinate Regression with Convolutional Neural Networks
- LIFT-SLAM: a deep-learning feature-based monocular visual SLAM method
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Recurrent Transformer Networks for Semantic Correspondence
- Large-Scale Image Retrieval with Attentive Deep Local Features
- Time-Contrastive Networks: Self-Supervised Learning from Video
- Toward Geometric Deep SLAM
- Self-Improving Visual Odometry
- Estimating 6D Pose From Localizing Designated Surface Keypoints
- SMPLR: Deep SMPL reverse for 3D human pose and shape recovery
- PPF-FoldNet: Unsupervised Learning of Rotation Invariant 3D Local Descriptors
- Integral Human Pose Regression
- FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence
- RegNet: Learning the Optimization of Direct Image-to-Image Pose Registration
- Local Feature Detectors, Descriptors, and Image Representations: A Survey
- Vision-based Pose Estimation for Augmented Reality : A Comparison Study
- DSAC - Differentiable RANSAC for Camera Localization
- ENG: End-to-end Neural Geometry for Robust Depth and Pose Estimation using CNNs
- Learning Spread-out Local Feature Descriptors
- Recurrent Neural Network for Learning DenseDepth and Ego-Motion from Video
- A Performance Evaluation of Local Features for Image Based 3D Reconstruction
- 3D Human Pose Estimation with 2D Marginal Heatmaps
- Image Matching: An Application-oriented Benchmark
- Multi-Image Semantic Matching by Mining Consistent Features
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- Quad-networks: unsupervised learning to rank for interest point detection
- Deep Spectral Correspondence for Matching Disparate Image Pairs
- Improving Nighttime Retrieval-Based Localization
- Learning Local Shape Descriptors from Part Correspondences With Multi-view Convolutional Networks
- REST: Real-to-Synthetic Transform for Illumination Invariant Camera Localization
- Specular-to-Diffuse Translation for Multi-View Reconstruction
- SConE: Siamese Constellation Embedding Descriptor for Image Matching
- Understanding and Improving Kernel Local Descriptors
- Learning to Align Images using Weak Geometric Supervision
- Matching Disparate Image Pairs Using Shape-Aware ConvNets