The Devil is in the Middle: Exploiting Mid-level Representations for Cross-Domain Instance Matching
arXiv:1711.08106
Abstract
Many vision problems require matching images of object instances across different domains. These include fine-grained sketch-based image retrieval (FG-SBIR) and Person Re-identification (person ReID). Existing approaches attempt to learn a joint embedding space where images from different domains can be directly compared. In most cases, this space is defined by the output of the final layer of a deep neural network (DNN), which primarily contains features of a high semantic level. In this paper, we argue that both high and mid-level features are relevant for cross-domain instance matching (CDIM). Importantly, mid-level features already exist in earlier layers of the DNN. They just need to be extracted, represented, and fused properly with the final layer. Based on this simple but powerful idea, we propose a unified framework for CDIM. Instantiating our framework for FG-SBIR and ReID, we show that our simple models can easily beat the state-of-the-art models, which are often equipped with much more elaborate architectures.
Reference updated
References in corpus (12)
- In Defense of the Triplet Loss for Person Re-Identification
- Improving Person Re-identification by Attribute and Identity Learning
- Deep Transfer Learning for Person Re-identification
- Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in vitro
- Show and Tell: A Neural Image Caption Generator
- Deeply-Learned Part-Aligned Representations for Person Re-Identification
- SVDNet for Pedestrian Retrieval
- Re-ranking Person Re-identification with k-reciprocal Encoding
- HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis
- Scalable Person Re-identification on Supervised Smoothed Manifold
- Person Re-Identification by Deep Joint Learning of Multi-Loss Classification
- Divide and Fuse: A Re-ranking Approach for Person Re-identification
Cited by in corpus (21)
- Omni-Scale Feature Learning for Person Re-Identification
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Incomplete Descriptor Mining with Elastic Loss for Person Re-Identification
- Parameter-Efficient Person Re-identification in the 3D Space
- Improved Person Re-Identification Based on Saliency and Semantic Parsing with Deep Neural Network Models
- Learning Generalisable Omni-Scale Representations for Person Re-Identification
- A Novel Teacher-Student Learning Framework For Occluded Person Re-Identification
- Sampling Agnostic Feature Representation for Long-Term Person Re-identification
- CA3Net: Contextual-Attentional Attribute-Appearance Network for Person Re-Identification
- CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification
- HAT: Hierarchical Aggregation Transformers for Person Re-identification
- Attention: A Big Surprise for Cross-Domain Person Re-Identification
- End-to-end Person Search Sequentially Trained on Aggregated Dataset
- Query Attack via Opposite-Direction Feature:Towards Robust Image Retrieval
- Multi-person Articulated Tracking with Spatial and Temporal Embeddings
- Long-Term Cloth-Changing Person Re-identification
- Connecting Language and Vision for Natural Language-Based Vehicle Retrieval
- STADB: A Self-Thresholding Attention Guided ADB Network for Person Re-identification
- HPILN: A feature learning framework for cross-modality person re-identification
- Copy and Paste method based on Pose for Re-identification
- Pose-Guided Feature Learning with Knowledge Distillation for Occluded Person Re-Identification