SCNet: Learning Semantic Correspondence
arXiv:1705.04043
Abstract
This paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearance only. We propose instead a convolutional neural network architecture, called SCNet, for learning a geometrically plausible model for semantic correspondence. SCNet uses region proposals as matching primitives, and explicitly incorporates geometric consistency in its loss function. It is trained on image pairs obtained from the PASCAL VOC 2007 keypoint dataset, and a comparative evaluation on several standard benchmarks demonstrates that the proposed approach substantially outperforms both recent deep learning architectures and previous methods based on hand-crafted features.
ICCV 2017
Cited by in corpus (9)
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
- Recurrent Transformer Networks for Semantic Correspondence
- CATs: Cost Aggregation Transformers for Visual Correspondence
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Joint Learning of Semantic Alignment and Object Landmark Detection
- Semantic Attribute Matching Networks
- iSPA-Net: Iterative Semantic Pose Alignment Network
- PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
- Object Pose Estimation from Monocular Image using Multi-View Keypoint Correspondence