Cross-Domain Visual Matching via Generalized Similarity Measure and Feature Learning
arXiv:1605.04039 · doi:10.1109/TPAMI.2016.2567386
Abstract
Cross-domain visual data matching is one of the fundamental problems in many real-world vision tasks, e.g., matching persons across ID photos and surveillance videos. Conventional approaches to this problem usually involves two steps: i) projecting samples from different domains into a common space, and ii) computing (dis-)similarity in this space based on a certain distance. In this paper, we present a novel pairwise similarity measure that advances existing models by i) expanding traditional linear projections into affine transformations and ii) fusing affine Mahalanobis distance and Cosine similarity by a data-driven combination. Moreover, we unify our similarity measure with feature representation learning via deep convolutional neural networks. Specifically, we incorporate the similarity measure matrix into the deep architecture, enabling an end-to-end way of model optimization. We extensively evaluate our generalized similarity model in several challenging cross-domain matching tasks: person re-identification under different views and face verification over different modalities (i.e., faces from still images and videos, older and younger faces, and sketch and photo portraits). The experimental results demonstrate superior performance of our model over other state-of-the-art methods.
To appear in IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2016
References in corpus (4)
Cited by in corpus (20)
- Deep Face Recognition: A Survey
- SCAN: Self-and-Collaborative Attention Network for Video Person Re-identification
- Content-Adaptive Sketch Portrait Generation by Decompositional Representation Learning
- An Efficient Framework for Visible-Infrared Cross Modality Person Re-Identification
- Instance-Aware Representation Learning and Association for Online Multi-Person Tracking
- M2M-GAN: Many-to-Many Generative Adversarial Transfer Learning for Person Re-Identification
- Hi-CMD: Hierarchical Cross-Modality Disentanglement for Visible-Infrared Person Re-Identification
- Style Normalization and Restitution for Generalizable Person Re-identification
- Attend to the Difference: Cross-Modality Person Re-identification via Contrastive Correlation
- Robust Metric Learning based on the Rescaled Hinge Loss
- Adaptively Connected Neural Networks
- Spatial-Temporal Person Re-identification
- Hetero-Center Loss for Cross-Modality Person Re-Identification
- Orthogonal Deep Features Decomposition for Age-Invariant Face Recognition
- Weakly Supervised Person Re-ID: Differentiable Graphical Learning and A New Benchmark
- Exploring Modality-shared Appearance Features and Modality-invariant Relation Features for Cross-modality Person Re-Identification
- Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning
- Unsupervised Multi-modal Hashing for Cross-modal retrieval
- Large Margin Learning in Set to Set Similarity Comparison for Person Re-identification
- Person Re-Identification using Deep Learning Networks: A Systematic Review