MegDet: A Large Mini-Batch Object Detector
arXiv:1711.07240
Abstract
The improvements in recent CNN-based object detection works, from R-CNN [11], Fast/Faster R-CNN [10, 31] to recent Mask R-CNN [14] and RetinaNet [24], mainly come from new network, new framework, or novel loss design. But mini-batch size, a key factor in the training, has not been well studied. In this paper, we propose a Large MiniBatch Object Detector (MegDet) to enable the training with much larger mini-batch size than before (e.g. from 16 to 256), so that we can effectively utilize multiple GPUs (up to 128 in our experiments) to significantly shorten the training time. Technically, we suggest a learning rate policy and Cross-GPU Batch Normalization, which together allow us to successfully train a large mini-batch detector in much less time (e.g., from 33 hours to 4 hours), and achieve even better accuracy. The MegDet is the backbone of our submission (mmAP 52.5%) to COCO 2017 Challenge, where we won the 1st place of Detection task.
References in corpus (8)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Focal Loss for Dense Object Detection
- cuDNN: Efficient Primitives for Deep Learning
- One weird trick for parallelizing convolutional neural networks
- YOLO9000: Better, Faster, Stronger
- Deformable Convolutional Networks
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- ImageNet Training in Minutes
Cited by in corpus (20)
- EfficientDet: Scalable and Efficient Object Detection
- Differentiable Learning-to-Normalize via Switchable Normalization
- Context Encoding for Semantic Segmentation
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Rethinking ImageNet Pre-training
- Retina U-Net: Embarrassingly Simple Exploitation of Segmentation Supervision for Medical Object Detection
- Scale-Aware Trident Networks for Object Detection
- One-Shot Instance Segmentation
- Soft Sampling for Robust Object Detection
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- Scene Text Detection with Supervised Pyramid Context Network
- UPSNet: A Unified Panoptic Segmentation Network
- Unified Perceptual Parsing for Scene Understanding
- Joint COCO and Mapillary Workshop at ICCV 2019: COCO Instance Segmentation Challenge Track
- Instance Shadow Detection
- Two-Stream Video Classification with Cross-Modality Attention
- G-RCN: Optimizing the Gap between Classification and Localization Tasks for Object Detection
- FA-RPN: Floating Region Proposals for Face Detection
- Quality-Aware Network for Face Parsing
- Efficient Scene Text Detection with Textual Attention Tower