Decoupled Classification Refinement: Hard False Positive Suppression for Object Detection
arXiv:1810.04002
Abstract
In this paper, we analyze failure cases of state-of-the-art detectors and observe that most hard false positives result from classification instead of localization and they have a large negative impact on the performance of object detectors. We conjecture there are three factors: (1) Shared feature representation is not optimal due to the mismatched goals of feature learning for classification and localization; (2) multi-task learning helps, yet optimization of the multi-task loss may result in sub-optimal for individual tasks; (3) large receptive field for different scales leads to redundant context information for small objects. We demonstrate the potential of detector classification power by a simple, effective, and widely-applicable Decoupled Classification Refinement (DCR) network. In particular, DCR places a separate classification network in parallel with the localization network (base detector). With ROI Pooling placed on the early stage of the classification network, we enforce an adaptive receptive field in DCR. During training, DCR samples hard false positives from the base detector and trains a strong classifier to refine classification results. During testing, DCR refines all boxes from the base detector. Experiments show competitive results on PASCAL VOC and COCO without any bells and whistles. Our codes are available at: https://github.com/bowenc0221/Decoupled-Classification-Refinement.
under review. arXiv admin note: substantial text overlap with arXiv:1803.06799
References in corpus (14)
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- DSSD : Deconvolutional Single Shot Detector
- Focal Loss for Dense Object Detection
- FCOS: Fully Convolutional One-Stage Object Detection
- Deformable Convolutional Networks
- CenterNet: Keypoint Triplets for Object Detection
- CornerNet: Detecting Objects as Paired Keypoints
- Bottom-up Object Detection by Grouping Extreme and Center Points
- Perceptual Generative Adversarial Networks for Small Object Detection
- Acquisition of Localization Confidence for Accurate Object Detection
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- SPGNet: Semantic Prediction Guidance for Scene Parsing
- Improving Object Detection from Scratch via Gated Feature Reuse
Cited by in corpus (14)
- Deep Learning for Generic Object Detection: A Survey
- SkyNet: a Hardware-Efficient Method for Object Detection and Tracking on Embedded Systems
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis
- Utilizing the Instability in Weakly Supervised Object Detection
- False Detection (Positives and Negatives) in Object Detection
- The 1st Tiny Object Detection Challenge:Methods and Results
- Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- High Frequency Residual Learning for Multi-Scale Image Classification
- A Simple Non-i.i.d. Sampling Approach for Efficient Training and Better Generalization
- A Global to Local Double Embedding Method for Multi-person Pose Estimation