RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
arXiv:1901.03353
Abstract
Recently two-stage detectors have surged ahead of single-shot detectors in the accuracy-vs-speed trade-off. Nevertheless single-shot detectors are immensely popular in embedded vision applications. This paper brings single-shot detectors up to the same level as current two-stage techniques. We do this by improving training for the state-of-the-art single-shot detector, RetinaNet, in three ways: integrating instance mask prediction for the first time, making the loss function adaptive and more stable, and including additional hard examples in training. We call the resulting augmented network RetinaMask. The detection component of RetinaMask has the same computational cost as the original RetinaNet, but is more accurate. COCO test-dev results are up to 41.4 mAP for RetinaMask-101 vs 39.1mAP for RetinaNet-101, while the runtime is the same during evaluation. Adding Group Normalization increases the performance of RetinaMask-101 to 41.7 mAP. Code is at:https://github.com/chengyangfu/retinamask
References in corpus (3)
Cited by in corpus (18)
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- EmbedMask: Embedding Coupling for One-stage Instance Segmentation
- Where are the Masks: Instance Segmentation with Image-level Supervision
- CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection
- SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation
- Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
- MaskFace: multi-task face and landmark detector
- Instance Segmentation with Point Supervision
- FGN: Fully Guided Network for Few-Shot Instance Segmentation
- CenterMask: single shot instance segmentation with point representation
- Learning Gaussian Maps for Dense Object Detection
- Mask Encoding for Single Shot Instance Segmentation
- SRF-GAN: Super-Resolved Feature GAN for Multi-Scale Representation
- Human-centric Relation Segmentation: Dataset and Solution
- PointINS: Point-based Instance Segmentation
- Sparsity-Inducing Optimal Control via Differential Dynamic Programming
- Boundary Distribution Estimation for Precise Object Detection
- VideoClick: Video Object Segmentation with a Single Click