Universal Correspondence Network
arXiv:1606.03558
Abstract
We present a deep learning framework for accurate visual correspondences and demonstrate its effectiveness for both geometric and semantic matching, spanning across rigid motions to intra-class shape or appearance variations. In contrast to previous CNN-based approaches that optimize a surrogate patch similarity objective, we use deep metric learning to directly learn a feature space that preserves either geometric or semantic similarity. Our fully convolutional architecture, along with a novel correspondence contrastive loss allows faster training by effective reuse of computations, accurate gradient computation through the use of thousands of examples per image pair and faster testing with feed forward passes for keypoints, instead of for typical patch similarity methods. We propose a convolutional spatial transformer to mimic patch normalization in traditional features like SIFT, which is shown to dramatically boost accuracy for semantic correspondences across intra-class shape variations. Extensive experiments on KITTI, PASCAL, and CUB-2011 datasets demonstrate the significant advantages of our features over prior works that use either hand-constructed or learned features.
To appear at NIPS 2016 as full oral presentation
References in corpus (4)
Cited by in corpus (45)
- R2D2: Repeatable and Reliable Detector and Descriptor
- SGUIE-Net: Semantic Attention Guided Underwater Image Enhancement with Multi-Scale Perception
- CalibNet: Geometrically Supervised Extrinsic Calibration using 3D Spatial Transformer Networks
- Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation
- A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence
- FECANet: Boosting Few-Shot Semantic Segmentation with Feature-Enhanced Context-Aware Network
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Neural Best-Buddies: Sparse Cross-Domain Correspondence
- Self-supervised Keypoint Correspondences for Multi-Person Pose Estimation and Tracking in Videos
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- Cross-domain Correspondence Learning for Exemplar-based Image Translation
- GCNv2: Efficient Correspondence Prediction for Real-Time SLAM
- Graph-Structured Visual Imitation
- Robust Optical Flow Estimation in Rainy Scenes
- All Graphs Lead to Rome: Learning Geometric and Cycle-Consistent Representations with Graph Convolutional Networks
- Learning Metrics from Teachers: Compact Networks for Image Embedding
- ResFPN: Residual Skip Connections in Multi-Resolution Feature Pyramid Networks for Accurate Dense Pixel Matching
- Semantic Correspondence via 2D-3D-2D Cycle
- Correspondence Networks with Adaptive Neighbourhood Consensus
- PixelNN: Example-based Image Synthesis
- Multi-Image Semantic Matching by Mining Consistent Features
- Extremely Dense Point Correspondences using a Learned Feature Descriptor
- Bi-level Feature Alignment for Versatile Image Translation and Manipulation
- RANSAC-Flow: generic two-stage image alignment
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- Template NeRF: Towards Modeling Dense Shape Correspondences from Category-Specific Object Images
- Inability of spatial transformations of CNN feature maps to support invariant recognition
- Understanding Pixel-level 2D Image Semantics with 3D Keypoint Knowledge Engine
- Semantic Attribute Matching Networks
- Unsupervised Metric Relocalization Using Transform Consistency Loss
- Combining Deep Learning and Verification for Precise Object Instance Detection
- iSPA-Net: Iterative Semantic Pose Alignment Network
- Statistical transformer networks: learning shape and appearance models via self supervision
- Semantic Matching by Weakly Supervised 2D Point Set Registration
- PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
- Semantic Correspondence: A Hierarchical Approach
- Proposal Flow: Semantic Correspondences from Object Proposals
- Robust Angular Local Descriptor Learning
- Informative and Representative Triplet Selection for Multilabel Remote Sensing Image Retrieval
- Deep Semantic Matching with Foreground Detection and Cycle-Consistency
- Real-time Halfway Domain Reconstruction of Motion and Geometry
- Object Pose Estimation from Monocular Image using Multi-View Keypoint Correspondence
- Fully Self-Supervised Class Awareness in Dense Object Descriptors
- Learning View and Target Invariant Visual Servoing for Navigation
- Semi-Supervised Semantic Matching