Pedestrian Alignment Network for Large-scale Person Re-identification
arXiv:1707.00408 · doi:10.1109/TCSVT.2018.2873599
Abstract
Person re-identification (person re-ID) is mostly viewed as an image retrieval problem. This task aims to search a query person in a large image pool. In practice, person re-ID usually adopts automatic detectors to obtain cropped pedestrian images. However, this process suffers from two types of detector errors: excessive background and part missing. Both errors deteriorate the quality of pedestrian alignment and may compromise pedestrian matching due to the position and scale variances. To address the misalignment problem, we propose that alignment can be learned from an identification procedure. We introduce the pedestrian alignment network (PAN) which allows discriminative embedding learning and pedestrian alignment without extra annotations. Our key observation is that when the convolutional neural network (CNN) learns to discriminate between different identities, the learned feature maps usually exhibit strong activations on the human body rather than the background. The proposed network thus takes advantage of this attention mechanism to adaptively locate and align pedestrians within a bounding box. Visual examples show that pedestrians are better aligned with PAN. Experiments on three large-scale re-ID datasets confirm that PAN improves the discriminative ability of the feature embeddings and yields competitive accuracy with the state-of-the-art methods.
References in corpus (7)
- In Defense of the Triplet Loss for Person Re-Identification
- Person Re-identification: Past, Present and Future
- Improving Person Re-identification by Attribute and Identity Learning
- Deep Transfer Learning for Person Re-identification
- Looking Beyond Appearances: Synthetic Training Data for Deep CNNs in Re-identification
- Pose Invariant Embedding for Deep Person Re-identification
- Learning Correspondence Structures for Person Re-identification
Cited by in corpus (85)
- In Defense of the Triplet Loss for Person Re-Identification
- Attention Mechanisms in Computer Vision: A Survey
- Random Erasing Data Augmentation
- Each Part Matters: Local Patterns Facilitate Cross-view Geo-localization
- A Transformer-Based Feature Segmentation and Region Alignment Method For UAV-View Geo-Localization
- MHSA-Net: Multi-Head Self-Attention Network for Occluded Person Re-Identification
- VehicleNet: Learning Robust Visual Representation for Vehicle Re-identification
- Beyond Part Models: Person Retrieval with Refined Part Pooling (and a Strong Convolutional Baseline)
- SVDNet for Pedestrian Retrieval
- Second-order Non-local Attention Networks for Person Re-identification
- Multi-task Learning with Coarse Priors for Robust Part-aware Person Re-identification
- Re-ID done right: towards good practices for person re-identification
- Incomplete Descriptor Mining with Elastic Loss for Person Re-Identification
- Deep Attention Aware Feature Learning for Person Re-Identification
- Joint Discriminative and Generative Learning for Person Re-identification
- Recognizing Partial Biometric Patterns
- Attention-Aware Compositional Network for Person Re-identification
- Parameter-Efficient Person Re-identification in the 3D Space
- Discriminative Feature Learning with Foreground Attention for Person Re-Identification
- ABD-Net: Attentive but Diverse Person Re-Identification
- Deep-Person: Learning Discriminative Deep Features for Person Re-Identification
- Horizontal Pyramid Matching for Person Re-identification
- Let Features Decide for Themselves: Feature Mask Network for Person Re-identification
- Foreground-aware Pyramid Reconstruction for Alignment-free Occluded Person Re-identification
- Auto-ReID: Searching for a Part-aware ConvNet for Person Re-Identification
- Relation-Aware Global Attention for Person Re-identification
- CDPM: Convolutional Deformable Part Models for Semantically Aligned Person Re-identification
- Camera Style Adaptation for Person Re-identification
- Pyramidal Person Re-IDentification via Multi-Loss Dynamic Training
- StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition
- GCT: Graph Co-Training for Semi-Supervised Few-Shot Learning
- Understanding Image Retrieval Re-Ranking: A Graph Neural Network Perspective
- Batch DropBlock Network for Person Re-identification and Beyond
- CA3Net: Contextual-Attentional Attribute-Appearance Network for Person Re-Identification
- Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification
- Cross-Resolution Adversarial Dual Network for Person Re-Identification and Beyond
- Densely Semantically Aligned Person Re-Identification
- Exploring Shape Embedding for Cloth-Changing Person Re-Identification via 2D-3D Correspondences
- Re-Identification with Consistent Attentive Siamese Networks
- Machine Learning in Artificial Intelligence: Towards a Common Understanding
- Dissecting Person Re-identification from the Viewpoint of Viewpoint
- A heterogeneous branch and multi-level classification network for person re-identification
- View Confusion Feature Learning for Person Re-identification
- Spatial-Temporal Person Re-identification
- STNReID : Deep Convolutional Networks with Pairwise Spatial Transformer Networks for Partial Person Re-identification
- Deep Co-attention based Comparators For Relative Representation Learning in Person Re-identification
- Grafted network for person re-identification
- An Evaluation of Deep CNN Baselines for Scene-Independent Person Re-Identification
- End-to-End Deep Kronecker-Product Matching for Person Re-identification
- Weighted Bilinear Coding over Salient Body Parts for Person Re-identification
- Hybrid-Attention Guided Network with Multiple Resolution Features for Person Re-Identification
- Exploring Modality-shared Appearance Features and Modality-invariant Relation Features for Cross-modality Person Re-Identification
- Learning to Learn in a Semi-Supervised Fashion
- Discovering Underlying Person Structure Pattern with Relative Local Distance for Person Re-identification
- Unsupervised Data Uncertainty Learning in Visual Retrieval Systems
- Pedestrian re-identification based on Tree branch network with local and global learning
- Deep Miner: A Deep and Multi-branch Network which Mines Rich and Diverse Features for Person Re-identification
- Exploring Uncertainty in Conditional Multi-Modal Retrieval Systems
- Collaborative Attention Network for Person Re-identification
- Unsupervised Eyeglasses Removal in the Wild
- Adaptive Re-ranking of Deep Feature for Person Re-identification
- Feature Affinity based Pseudo Labeling for Semi-supervised Person Re-identification
- Integrating Coarse Granularity Part-level Features with Supervised Global-level Features for Person Re-identification
- Inability of spatial transformations of CNN feature maps to support invariant recognition
- Improved Res2Net model for Person re-identification
- In Defense of the Classification Loss for Person Re-Identification
- Multigranular Visual-Semantic Embedding for Cloth-Changing Person Re-identification
- Homocentric Hypersphere Feature Embedding for Person Re-identification
- Hierarchical and Efficient Learning for Person Re-Identification
- Progressive Multi-stage Feature Mix for Person Re-Identification
- GAN-based Pose-aware Regulation for Video-based Person Re-identification
- Learning Deep Representations by Mutual Information for Person Re-identification
- Person Re-Identification using Deep Learning Networks: A Systematic Review
- Virtual CNN Branching: Efficient Feature Ensemble for Person Re-Identification
- Video-based Person Re-identification without Bells and Whistles
- Single Camera Training for Person Re-identification
- Weakly Supervised Tracklet Person Re-Identification by Deep Feature-wise Mutual Learning
- Person Re-identification with Adversarial Triplet Embedding
- Cross-Resolution Person Re-identification with Deep Antithetical Learning
- Beyond Triplet Loss: Meta Prototypical N-tuple Loss for Person Re-identification
- MDFM: Multi-Decision Fusing Model for Few-Shot Learning
- Pose Invariant Person Re-Identification using Robust Pose-transformation GAN
- Attribute Guided Sparse Tensor-Based Model for Person Re-Identification
- VMRFANet:View-Specific Multi-Receptive Field Attention Network for Person Re-identification
- Tasks Integrated Networks: Joint Detection and Retrieval for Image Search