HA-CCN: Hierarchical Attention-based Crowd Counting Network
arXiv:1907.10255 · doi:10.1109/TIP.2019.2928634
Abstract
Single image-based crowd counting has recently witnessed increased focus, but many leading methods are far from optimal, especially in highly congested scenes. In this paper, we present Hierarchical Attention-based Crowd Counting Network (HA-CCN) that employs attention mechanisms at various levels to selectively enhance the features of the network. The proposed method, which is based on the VGG16 network, consists of a spatial attention module (SAM) and a set of global attention modules (GAM). SAM enhances low-level features in the network by infusing spatial segmentation information, whereas the GAM focuses on enhancing channel-wise information in the higher level layers. The proposed method is a single-step training framework, simple to implement and achieves state-of-the-art results on different datasets. Furthermore, we extend the proposed counting network by introducing a novel set-up to adapt the network to different scenes and datasets via weak supervision using image-level labels. This new set up reduces the burden of acquiring labour intensive point-wise annotations for new datasets while improving the cross-dataset performance.
Accepted for publication at IEEE Transactions on Image Processing (TIP) 2019
References in corpus (4)
Cited by in corpus (13)
- A General Survey on Attention Mechanisms in Deep Learning
- A Self-Training Approach for Point-Supervised Object Detection and Counting in Crowds
- Neuron Linear Transformation: Modeling the Domain Shift for Crowd Counting
- Encoder-Decoder Based Convolutional Neural Networks with Multi-Scale-Aware Modules for Crowd Counting
- Spatiotemporal Dilated Convolution with Uncertain Matching for Video-based Crowd Estimation
- Crowd Counting via Perspective-Guided Fractional-Dilation Convolution
- Dilated-Scale-Aware Attention ConvNet For Multi-Class Object Counting
- TreeFormer: a Semi-Supervised Transformer-based Framework for Tree Counting from a Single High Resolution Image
- Video Crowd Localization with Multi-focus Gaussian Neighborhood Attention and a Large-Scale Benchmark
- Forget Less, Count Better: A Domain-Incremental Self-Distillation Learning Benchmark for Lifelong Crowd Counting
- Learning Discriminative Features for Crowd Counting
- BBA-net: A bi-branch attention network for crowd counting
- An Improved Normed-Deformable Convolution for Crowd Counting