What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
arXiv:1707.07074 · doi:10.1016/j.patcog.2017.10.004
Abstract
Matching pedestrians across disjoint camera views, known as person re-identification (re-id), is a challenging problem that is of importance to visual recognition and surveillance. Most existing methods exploit local regions within spatial manipulation to perform matching in local correspondence. However, they essentially extract \emph{fixed} representations from pre-divided regions for each image and perform matching based on the extracted representation subsequently. For models in this pipeline, local finer patterns that are crucial to distinguish positive pairs from negative ones cannot be captured, and thus making them underperformed. In this paper, we propose a novel deep multiplicative integration gating function, which answers the question of \emph{what-and-where to match} for effective person re-id. To address \emph{what} to match, our deep network emphasizes common local patterns by learning joint representations in a multiplicative way. The network comprises two Convolutional Neural Networks (CNNs) to extract convolutional activations, and generates relevant descriptors for pedestrian matching. This thus, leads to flexible representations for pair-wise images. To address \emph{where} to match, we combat the spatial misalignment by performing spatially recurrent pooling via a four-directional recurrent neural network to impose spatial dependency over all positions with respect to the entire image. The proposed network is designed to be end-to-end trainable to characterize local pairwise feature interactions in a spatially aligned manner. To demonstrate the superiority of our method, extensive experiments are conducted over three benchmark data sets: VIPeR, CUHK03 and Market-1501.
Published at Pattern Recognition, Elsevier
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Person Re-Identification by Camera Correlation Aware Feature Augmentation
- Deep Adaptive Feature Embedding with Local Sample Distributions for Person Re-identification
- Pose Invariant Embedding for Deep Person Re-identification
- Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
- Embedding Deep Metric for Person Re-identication A Study Against Large Variations
- A Siamese Long Short-Term Memory Architecture for Human Re-Identification
- Scalable Person Re-identification on Supervised Smoothed Manifold
Cited by in corpus (17)
- Multi-View Spectral Clustering via Structured Low-Rank Matrix Factorization
- Multi-Domain Adversarial Feature Generalization for Person Re-Identification
- Attribute-guided Feature Learning Network for Vehicle Re-identification
- Where to Focus: Deep Attention-based Spatially Recurrent Bilinear Networks for Fine-Grained Visual Recognition
- Cross-Entropy Adversarial View Adaptation for Person Re-identification
- Eliminating cross-camera bias for vehicle re-identification
- Cross Domain Knowledge Learning with Dual-branch Adversarial Network for Vehicle Re-identification
- Deep neural network-based classification model for Sentiment Analysis
- Person Re-Identification using Deep Learning Networks: A Systematic Review
- PAC-GAN: An Effective Pose Augmentation Scheme for Unsupervised Cross-View Person Re-identification
- Hierarchical Attention Network for Action Segmentation
- A Targeted Acceleration and Compression Framework for Low bit Neural Networks
- Using Context Information to Enhance Simple Question Answering
- Anomaly detecting and ranking of the cloud computing platform by multi-view learning
- Multi-feature Distance Metric Learning for Non-rigid 3D Shape Retrieval
- Auto-weighted Mutli-view Sparse Reconstructive Embedding
- Structured Mean-field Variational Inference and Learning in Winner-take-all Spiking Neural Networks