End-to-end people detection in crowded scenes
arXiv:1506.04878
Abstract
Current people detectors operate either by scanning an image in a sliding window fashion or by classifying a discrete set of proposals. We propose a model that is based on decoding an image into a set of people detections. Our system takes an image as input and directly outputs a set of distinct detection hypotheses. Because we generate predictions jointly, common post-processing steps such as non-maximum suppression are unnecessary. We use a recurrent LSTM layer for sequence generation and train our model end-to-end with a new loss function that operates on sets of detections. We demonstrate the effectiveness of our approach on the challenging task of detecting people in crowded scenes.
9 pages, 7 figures. Submitted to NIPS 2015. Supplementary material video: http://www.youtube.com/watch?v=QeWl0h3kQ24
References in corpus (4)
Cited by in corpus (7)
- Semantic Instance Segmentation with a Discriminative Loss Function
- Switching Convolutional Neural Network for Crowd Counting
- A Computer Vision System to Localize and Classify Wastes on the Streets
- Learning non-maximum suppression
- Recurrent Filter Learning for Visual Tracking
- Sequential Person Recognition in Photo Albums with a Recurrent Network
- A Video Analysis Method on Wanfang Dataset via Deep Neural Network