Accurate Single Stage Detector Using Recurrent Rolling Convolution
arXiv:1704.05776
Abstract
Most of the recent successful methods in accurate object detection and localization used some variants of R-CNN style two stage Convolutional Neural Networks (CNN) where plausible regions were proposed in the first stage then followed by a second stage for decision refinement. Despite the simplicity of training and the efficiency in deployment, the single stage detection methods have not been as competitive when evaluated in benchmarks consider mAP for high IoU thresholds. In this paper, we proposed a novel single stage end-to-end trainable object detection network to overcome this limitation. We achieved this by introducing Recurrent Rolling Convolution (RRC) architecture over multi-scale feature maps to construct object classifiers and bounding box regressors which are "deep in context". We evaluated our method in the challenging KITTI dataset which measures methods under IoU threshold of 0.7. We showed that with RRC, a single reduced VGG-16 based model already significantly outperformed all the previously published results. At the time this paper was written our models ranked the first in KITTI car detection (the hard level), the first in cyclist detection and the second in pedestrian detection. These results were not reached by the previous single stage methods. The code is publicly available.
CVPR 2017
References in corpus (2)
Cited by in corpus (19)
- Track to Reconstruct and Reconstruct to Track
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- A Survey on Deep Learning Methods for Robot Vision
- DeepSignals: Predicting Intent of Drivers Through Visual Signals
- Residual Features and Unified Prediction Network for Single Stage Detection
- Extend the shallow part of Single Shot MultiBox Detector via Convolutional Neural Network
- Fusing Bird View LIDAR Point Cloud and Front View Camera Image for Deep Object Detection
- Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds
- MDFN: Multi-Scale Deep Feature Learning Network for Object Detection
- Weaving Multi-scale Context for Single Shot Detector
- Exploring Multi-Branch and High-Level Semantic Networks for Improving Pedestrian Detection
- On Machine Learning and Structure for Mobile Robots
- MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object Detection
- Relationship-Embedded Representation Learning for Grounding Referring Expressions
- Pedestrian Detection with Autoregressive Network Phases
- Multiple receptive fields and small-object-focusing weakly-supervised segmentation network for fast object detection
- Pseudo-labels for Supervised Learning on Dynamic Vision Sensor Data, Applied to Object Detection under Ego-motion
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- Content-adaptive Representation Learning for Fast Image Super-resolution