Learning Generalisable Omni-Scale Representations for Person Re-Identification
arXiv:1910.06827
Abstract
An effective person re-identification (re-ID) model should learn feature representations that are both discriminative, for distinguishing similar-looking people, and generalisable, for deployment across datasets without any adaptation. In this paper, we develop novel CNN architectures to address both challenges. First, we present a re-ID CNN termed omni-scale network (OSNet) to learn features that not only capture different spatial scales but also encapsulate a synergistic combination of multiple scales, namely omni-scale features. The basic building block consists of multiple convolutional streams, each detecting features at a certain scale. For omni-scale feature learning, a unified aggregation gate is introduced to dynamically fuse multi-scale features with channel-wise weights. OSNet is lightweight as its building blocks comprise factorised convolutions. Second, to improve generalisable feature learning, we introduce instance normalisation (IN) layers into OSNet to cope with cross-dataset discrepancies. Further, to determine the optimal placements of these IN layers in the architecture, we formulate an efficient differentiable architecture search algorithm. Extensive experiments show that, in the conventional same-dataset setting, OSNet achieves state-of-the-art performance, despite being much smaller than existing re-ID models. In the more challenging yet practical cross-dataset setting, OSNet beats most recent unsupervised domain adaptation methods without using any target data. Our code and models are released at \texttt{https://github.com/KaiyangZhou/deep-person-reid}.
TPAMI 2021. Journal extension of arXiv:1905.00953. Updates: added appendix. arXiv admin note: text overlap with arXiv:1905.00953
References in corpus (11)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Neural Architecture Search with Reinforcement Learning
- On the Convergence of Adam and Beyond
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- A Learned Representation For Artistic Style
- Random Erasing Data Augmentation
- GLAD: Global-Local-Alignment Descriptor for Pedestrian Retrieval
- Deep Transfer Learning for Person Re-identification
- Monte Carlo Gradient Estimation in Machine Learning
- Domain Generalization with MixStyle
Cited by in corpus (13)
- Omni-Scale Feature Learning for Person Re-Identification
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Learning to Generate Novel Domains for Domain Generalization
- MixStyle Neural Networks for Domain Generalization and Adaptation
- Rank Flow Embedding for Unsupervised and Semi-Supervised Manifold Learning
- Building Computationally Efficient and Well-Generalizing Person Re-Identification Models with Metric Learning
- Meta Batch-Instance Normalization for Generalizable Person Re-Identification
- Deep Domain-Adversarial Image Generation for Domain Generalisation
- Unsupervised Disentanglement GAN for Domain Adaptive Person Re-Identification
- Symbiotic Adversarial Learning for Attribute-based Person Search
- A Strong Baseline for Fashion Retrieval with Person Re-Identification Models
- Channel Recurrent Attention Networks for Video Pedestrian Retrieval
- Psychophysical Evaluation of Deep Re-Identification Models