TransCrowd: weakly-supervised crowd counting with transformers
arXiv:2104.09116 · doi:10.1007/s11432-021-3445-y
Abstract
The mainstream crowd counting methods usually utilize the convolution neural network (CNN) to regress a density map, requiring point-level annotations. However, annotating each person with a point is an expensive and laborious process. During the testing phase, the point-level annotations are not considered to evaluate the counting accuracy, which means the point-level annotations are redundant. Hence, it is desirable to develop weakly-supervised counting methods that just rely on count-level annotations, a more economical way of labeling. Current weakly-supervised counting methods adopt the CNN to regress a total count of the crowd by an image-to-count paradigm. However, having limited receptive fields for context modeling is an intrinsic limitation of these weakly-supervised CNN-based methods. These methods thus cannot achieve satisfactory performance, with limited applications in the real world. The transformer is a popular sequence-to-sequence prediction model in natural language processing (NLP), which contains a global receptive field. In this paper, we propose TransCrowd, which reformulates the weakly-supervised crowd counting problem from the perspective of sequence-to-count based on transformers. We observe that the proposed TransCrowd can effectively extract the semantic crowd information by using the self-attention mechanism of transformer. To the best of our knowledge, this is the first work to adopt a pure transformer for crowd counting research. Experiments on five benchmark datasets demonstrate that the proposed TransCrowd achieves superior performance compared with all the weakly-supervised CNN-based counting methods and gains highly competitive counting performance compared with some popular fully-supervised counting methods.
Accepted by Science China Information Sciences (SCIS). Code is available at https://github.com/dk-liang/TransCrowd
References in corpus (5)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Exploiting Unlabeled Data in CNNs by Self-supervised Learning to Rank
- Focal Inverse Distance Transform Maps for Crowd Localization
- Dilated-Scale-Aware Attention ConvNet For Multi-Class Object Counting
Cited by in corpus (9)
- End-to-end Temporal Action Detection with Transformer
- Focal Inverse Distance Transform Maps for Crowd Localization
- Counting Varying Density Crowds Through Density Guided Adaptive Selection CNN and Transformer Estimation
- TreeFormer: a Semi-Supervised Transformer-based Framework for Tree Counting from a Single High Resolution Image
- CrowdFormer: Weakly-supervised Crowd counting with Improved Generalizability
- Density-based clustering with fully-convolutional networks for crowd flow detection from drones
- STB-VMM: Swin Transformer Based Video Motion Magnification
- Learning Discriminative Features for Crowd Counting
- Counting Manatee Aggregations using Deep Neural Networks and Anisotropic Gaussian Kernel