Style Aggregated Network for Facial Landmark Detection
arXiv:1803.04108
Abstract
Recent advances in facial landmark detection achieve success by learning discriminative features from rich deformation of face shapes and poses. Besides the variance of faces themselves, the intrinsic variance of image styles, e.g., grayscale vs. color images, light vs. dark, intense vs. dull, and so on, has constantly been overlooked. This issue becomes inevitable as increasing web images are collected from various sources for training neural networks. In this work, we propose a style-aggregated approach to deal with the large intrinsic variance of image styles for facial landmark detection. Our method transforms original face images to style-aggregated images by a generative adversarial module. The proposed scheme uses the style-aggregated image to maintain face images that are more robust to environmental changes. Then the original face images accompanying with style-aggregated ones play a duet to train a landmark detector which is complementary to each other. In this way, for each face, our method takes two images as input, i.e., one in its original style and the other in the aggregated style. In experiments, we observe that the large variance of image styles would degenerate the performance of facial landmark detectors. Moreover, we show the robustness of our method to the large variance of image styles by comparing to a variant of our approach, in which the generative adversarial module is removed, and no style-aggregated images are used. Our approach is demonstrated to perform well when compared with state-of-the-art algorithms on benchmark datasets AFLW and 300-W. Code is publicly available on GitHub: https://github.com/D-X-Y/SAN
Accepted to CVPR 2018
References in corpus (14)
- R-FCN: Object Detection via Region-based Fully Convolutional Networks
- InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets
- A Learned Representation For Artistic Style
- f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization
- Texture Networks: Feed-forward Synthesis of Textures and Stylized Images
- Stacked Hourglass Networks for Human Pose Estimation
- Convolutional Pose Machines
- LR-GAN: Layered Recursive Generative Adversarial Networks for Image Generation
- AdaGAN: Boosting Generative Models
- Face Aging With Conditional Generative Adversarial Networks
- More is Less: A More Complicated Network with Less Inference Complexity
- A Recurrent Encoder-Decoder Network for Sequential Face Alignment
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- Pose-Invariant Face Alignment with a Single CNN
Cited by in corpus (17)
- Camera Style Adaptation for Person Re-identification
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- Look at Boundary: A Boundary-Aware Face Alignment Algorithm
- LOTR: Face Landmark Localization Using Localization Transformer
- ReenactGAN: Learning to Reenact Faces via Boundary Transfer
- Progressive Sample Mining and Representation Learning for One-Shot Person Re-identification with Adversarial Samples
- Deep Adversarial Attention Alignment for Unsupervised Domain Adaptation: the Benefit of Target Expectation Maximization
- Teacher Supervises Students How to Learn From Partially Labeled Images for Facial Landmark Detection
- Adaloss: Adaptive Loss Function for Landmark Localization
- Beyond Trade-off: Accelerate FCN-based Face Detector with Higher Accuracy
- 2D Wasserstein Loss for Robust Facial Landmark Detection
- KPNet: Towards Minimal Face Detector
- ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment
- FineNet: Frame Interpolation and Enhancement for Face Video Deblurring
- Think about boundary: Fusing multi-level boundary information for landmark heatmap regression
- Robust Facial Landmark Detection by Cross-order Cross-semantic Deep Network
- Human Recognition Using Face in Computed Tomography