Learning Discriminative Features for Crowd Counting
arXiv:2311.04509 · doi:10.1109/TIP.2024.3408609
Abstract
Crowd counting models in highly congested areas confront two main challenges: weak localization ability and difficulty in differentiating between foreground and background, leading to inaccurate estimations. The reason is that objects in highly congested areas are normally small and high level features extracted by convolutional neural networks are less discriminative to represent small objects. To address these problems, we propose a learning discriminative features framework for crowd counting, which is composed of a masked feature prediction module (MPM) and a supervised pixel-level contrastive learning module (CLM). The MPM randomly masks feature vectors in the feature map and then reconstructs them, allowing the model to learn about what is present in the masked regions and improving the model's ability to localize objects in high density regions. The CLM pulls targets close to each other and pushes them far away from background in the feature space, enabling the model to discriminate foreground objects from background. Additionally, the proposed modules can be beneficial in various computer vision tasks, such as crowd counting and object detection, where dense scenes or cluttered environments pose challenges to accurate localization. The proposed two modules are plug-and-play, incorporating the proposed modules into existing models can potentially boost their performance in these scenarios.
References in corpus (15)
- Adam: A Method for Stochastic Optimization
- NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting and Localization
- TransCrowd: weakly-supervised crowd counting with transformers
- HA-CCN: Hierarchical Attention-based Crowd Counting Network
- Focal Inverse Distance Transform Maps for Crowd Localization
- A Self-Training Approach for Point-Supervised Object Detection and Counting in Crowds
- PaDNet: Pan-Density Crowd Counting
- Redesigning Multi-Scale Neural Network for Crowd Counting
- Tracking-by-Counting: Using Network Flows on Crowd Density Maps for Tracking Multiple Targets
- Encoder-Decoder Based Convolutional Neural Networks with Multi-Scale-Aware Modules for Crowd Counting
- Counting Varying Density Crowds Through Density Guided Adaptive Selection CNN and Transformer Estimation
- Fine-Grained Crowd Counting
- FusionCount: Efficient Crowd Counting via Multiscale Feature Fusion
- Learning Independent Instance Maps for Crowd Localization
- Crowd Localization from Gaussian Mixture Scoped Knowledge and Scoped Teacher