Face Detection with Feature Pyramids and Landmarks
arXiv:1912.00596
Abstract
Accurate face detection and facial landmark localization are crucial to any face recognition system. We present a series of three single-stage RCNNs with different sized backbones (MobileNetV2-25, MobileNetV2-100, and ResNet101) and a six-layer feature pyramid trained exclusively on the WIDER FACE dataset. We compare the face detection and landmark accuracies using eight context module architectures, four proposed by previous research and four modified versions. We find no evidence that any of the proposed architectures significantly overperform and postulate that the random initialization of the additional layers is at least of equal importance. To show this we present a model that achieves near state-of-the-art performance on WIDER FACE and also provides high accuracy landmarks with a simple context module. We also present results using MobileNetV2 backbones, which achieve over 90% average precision on the WIDER FACE hard validation set while being able to run in real-time. By comparing to other authors, we show that our models exceed the state-of-the-art for similar-sized RCNNs and match the performance of much heavier networks.
12 pages, 2 figures, whitepaper
References in corpus (9)
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- Light-Head R-CNN: In Defense of Two-Stage Object Detector
- Face Attention Network: An Effective Face Detector for the Occluded Faces
- PyramidBox++: High Performance Detector for Finding Tiny Face
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- Improved Selective Refinement Network for Face Detection
- Accurate Face Detection for High Performance
- Robust and High Performance Face Detector
- Supervised Transformer Network for Efficient Face Detection