PN-Net: Conjoined Triple Deep Network for Learning Local Image Descriptors
arXiv:1601.05030
Abstract
In this paper we propose a new approach for learning local descriptors for matching image patches. It has recently been demonstrated that descriptors based on convolutional neural networks (CNN) can significantly improve the matching performance. Unfortunately their computational complexity is prohibitive for any practical application. We address this problem and propose a CNN based descriptor with improved matching performance, significantly reduced training and execution time, as well as low dimensionality. We propose to train the network with triplets of patches that include a positive and negative pairs. To that end we introduce a new loss function that exploits the relations within the triplets. We compare our approach to recently introduced MatchNet and DeepCompare and demonstrate the advantages of our descriptor in terms of performance, memory footprint and speed i.e. when run in GPU, the extraction time of our 128 dimensional feature is comparable to the fastest available binary descriptors such as BRIEF and ORB.
References in corpus (2)
Cited by in corpus (30)
- R2D2: Repeatable and Reliable Detector and Descriptor
- A Survey on Deep Learning Techniques for Stereo-based Depth Estimation
- GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints
- LIFT: Learned Invariant Feature Transform
- Generic 3D Representation via Pose Estimation and Matching
- UnsuperPoint: End-to-end Unsupervised Interest Point Detector and Descriptor
- Convolutional neural network architecture for geometric matching
- Deep Learning a Grasp Function for Grasping under Gripper Pose Uncertainty
- UR2KiD: Unifying Retrieval, Keypoint Detection, and Keypoint Description without Local Correspondence Supervision
- Explaining Away Results in Accurate and Tolerant Template Matching
- ContextDesc: Local Descriptor Augmentation with Cross-Modality Context
- Unsupervised learning from videos using temporal coherency deep networks
- From handcrafted to deep local features
- Deep unsupervised learning through spatial contrasting
- Bias and Fairness in Computer Vision Applications of the Criminal Justice System
- Hard-Aware Deeply Cascaded Embedding
- S2DNet: Learning Accurate Correspondences for Sparse-to-Dense Feature Matching
- Quasiconformal model with CNN features for large deformation image registration
- Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolutions
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- TS-Net: Combining modality specific and common features for multimodal patch matching
- The CUDA LATCH Binary Descriptor: Because Sometimes Faster Means Better
- Embedded Spectral Descriptors: Learning the point-wise correspondence metric via Siamese neural networks
- A Structure Feature Algorithm for Multi-modal Forearm Registration
- SDC - Stacked Dilated Convolution: A Unified Descriptor Network for Dense Matching Tasks
- Utilizing Complex-valued Network for Learning to Compare Image Patches
- mdBrief - A Fast Online Adaptable, Distorted Binary Descriptor for Real-Time Applications Using Calibrated Wide-Angle Or Fisheye Cameras
- Understanding and Improving Kernel Local Descriptors
- Learning to Align Images using Weak Geometric Supervision
- Semi-supervised learning of deep metrics for stereo reconstruction