LF-Net: Learning Local Features from Images
arXiv:1805.09662
Abstract
We present a novel deep architecture and a training strategy to learn a local feature pipeline from scratch, using collections of images without the need for human supervision. To do so we exploit depth and relative camera pose cues to create a virtual target that the network should achieve on one image, provided the outputs of the network for the other image. While this process is inherently non-differentiable, we show that we can optimize the network in a two-branch setup by confining it to one branch, while preserving differentiability in the other. We train our method on both indoor and outdoor datasets, with depth data from 3D sensors for the former, and depth estimates from an off-the-shelf Structure-from-Motion solution for the latter. Our models outperform the state of the art on sparse feature matching on both datasets, while running at 60+ fps for QVGA images.
NIPS 2018
Cited by in corpus (37)
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- DISK: Learning local features with policy gradient
- Reference Pose Generation for Long-term Visual Localization via Learned Features and View Synthesis
- Key.Net: Keypoint Detection by Handcrafted and Learned CNN Filters
- TopicFM: Robust and Interpretable Topic-Assisted Feature Matching
- Learning Two-View Correspondences and Geometry Using Order-Aware Network
- Wide-Baseline Relative Camera Pose Estimation with Directional Learning
- Unifying Deep Local and Global Features for Image Search
- Self-Improving Visual Odometry
- D3Feat: Joint Learning of Dense Detection and Description of 3D Local Features
- Probabilistic Spatial Distribution Prior Based Attentional Keypoints Matching Network
- HyNet: Learning Local Descriptor with Hybrid Similarity Measure and Triplet Loss
- Learning Second-Order Attentive Context for Efficient Correspondence Pruning
- ACNe: Attentive Context Normalization for Robust Permutation-Equivariant Learning
- Spatially Consistent Representation Learning
- Segmentation-driven 6D Object Pose Estimation
- Evaluation of Point Pattern Features for Anomaly Detection of Defect within Random Finite Set Framework
- COTR: Correspondence Transformer for Matching Across Images
- Self-Supervised 3D Keypoint Learning for Ego-motion Estimation
- StickyPillars: Robust and Efficient Feature Matching on Point Clouds using Graph Neural Networks
- Unsupervised Object-Level Representation Learning from Scene Images
- Learning Geodesic-Aware Local Features from RGB-D Images
- IV-SLAM: Introspective Vision for Simultaneous Localization and Mapping
- Classic versus deep learning approaches to address computer vision challenges
- Optimizing Through Learned Errors for Accurate Sports Field Registration
- Learning Feature Descriptors using Camera Pose Supervision
- Single-Stage 6D Object Pose Estimation
- Extracting Deformation-Aware Local Features by Learning to Deform
- Soft Expectation and Deep Maximization for Image Feature Detection
- Feature matching in Ultrasound images
- Linearized Multi-Sampling for Differentiable Image Transformation
- Unsupervised Learning Framework of Interest Point Via Properties Optimization
- ViewSynth: Learning Local Features from Depth using View Synthesis
- Large Scale Indexing of Generic Medical Image Data using Unbiased Shallow Keypoints and Deep CNN Features
- Learning Local Feature Descriptor with Motion Attribute for Vision-based Localization
- Asynchronous Multi-View SLAM
- Compression of descriptor models for mobile applications