Gradient Harmonized Single-stage Detector
arXiv:1811.05181
Abstract
Despite the great success of two-stage detectors, single-stage detector is still a more elegant and efficient way, yet suffers from the two well-known disharmonies during training, i.e. the huge difference in quantity between positive and negative examples as well as between easy and hard examples. In this work, we first point out that the essential effect of the two disharmonies can be summarized in term of the gradient. Further, we propose a novel gradient harmonizing mechanism (GHM) to be a hedging for the disharmonies. The philosophy behind GHM can be easily embedded into both classification loss function like cross-entropy (CE) and regression loss function like smooth- () loss. To this end, two novel loss functions called GHM-C and GHM-R are designed to balancing the gradient flow for anchor classification and bounding box refinement, respectively. Ablation study on MS COCO demonstrates that without laborious hyper-parameter tuning, both GHM-C and GHM-R can bring substantial improvement for single-stage detector. Without any whistles and bells, our model achieves 41.6 mAP on COCO test-dev set which surpasses the state-of-the-art method, Focal Loss (FL) + , by 0.8.
To appear at AAAI 2019
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- YOLOv3: An Incremental Improvement
- DSSD : Deconvolutional Single Shot Detector
- Focal Loss for Dense Object Detection
- GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
- Improving Regression Performance with Distributional Losses
Cited by in corpus (14)
- Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
- Convolutional Neural Networks with Gated Recurrent Connections
- Deep Representation Learning on Long-tailed Data: A Learnable Embedding Augmentation Perspective
- Reinterpreting CTC training as iterative fitting
- What is Next when Sequential Prediction Meets Implicitly Hard Interaction?
- KTN: Knowledge Transfer Network for Learning Multi-person 2D-3D Correspondences
- Feature Intertwiner for Object Detection
- AM-LFS: AutoML for Loss Function Search
- Rethinking Classification and Localization for Cascade R-CNN
- Learning from Noisy Anchors for One-stage Object Detection
- Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
- FPCD: An Open Aerial VHR Dataset for Farm Pond Change Detection
- Prime-Aware Adaptive Distillation
- ProbaNet: Proposal-balanced Network for Object Detection