CA3Net: Contextual-Attentional Attribute-Appearance Network for Person Re-Identification
arXiv:1811.07544
Abstract
Person re-identification aims to identify the same pedestrian across non-overlapping camera views. Deep learning techniques have been applied for person re-identification recently, towards learning representation of pedestrian appearance. This paper presents a novel Contextual-Attentional Attribute-Appearance Network (CA3Net) for person re-identification. The CA3Net simultaneously exploits the complementarity between semantic attributes and visual appearance, the semantic context among attributes, visual attention on attributes as well as spatial dependencies among body parts, leading to discriminative and robust pedestrian representation. Specifically, an attribute network within CA3Net is designed with an Attention-LSTM module. It concentrates the network on latent image regions related to each attribute as well as exploits the semantic context among attributes by a LSTM module. An appearance network is developed to learn appearance features from the full body, horizontal and vertical body parts of pedestrians with spatial dependencies among body parts. The CA3Net jointly learns the attribute and appearance features in a multi-task learning manner, generating comprehensive representation of pedestrians. Extensive experiments on two challenging benchmarks, i.e., Market-1501 and DukeMTMC-reID datasets, have demonstrated the effectiveness of the proposed approach.
References in corpus (17)
- In Defense of the Triplet Loss for Person Re-Identification
- Improving Person Re-identification by Attribute and Identity Learning
- Random Erasing Data Augmentation
- Pedestrian Alignment Network for Large-scale Person Re-identification
- GLAD: Global-Local-Alignment Descriptor for Pedestrian Retrieval
- Person Re-Identification by Camera Correlation Aware Feature Augmentation
- Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in vitro
- Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking
- Harmonious Attention Network for Person Re-Identification
- SVDNet for Pedestrian Retrieval
- Re-ranking Person Re-identification with k-reciprocal Encoding
- HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis
- Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification
- The Devil is in the Middle: Exploiting Mid-level Representations for Cross-Domain Instance Matching
- Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
- Deep-Person: Learning Discriminative Deep Features for Person Re-Identification
- Attribute Recognition by Joint Recurrent Learning of Context and Correlation
Cited by in corpus (6)
- Multi-scale 3D Convolution Network for Video Based Person Re-Identification
- Real-world Person Re-Identification via Degradation Invariance Learning
- Attention: A Big Surprise for Cross-Domain Person Re-Identification
- AttKGCN: Attribute Knowledge Graph Convolutional Network for Person Re-identification
- Temporal Attribute-Appearance Learning Network for Video-based Person Re-Identification
- Joint Discriminative and Metric Embedding Learning for Person Re-Identification