How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)
arXiv:1703.07332 · doi:10.1109/ICCV.2017.116
Abstract
This paper investigates how far a very deep neural network is from attaining close to saturating performance on existing 2D and 3D face alignment datasets. To this end, we make the following 5 contributions: (a) we construct, for the first time, a very strong baseline by combining a state-of-the-art architecture for landmark localization with a state-of-the-art residual block, train it on a very large yet synthetically expanded 2D facial landmark dataset and finally evaluate it on all other 2D facial landmark datasets. (b) We create a guided by 2D landmarks network which converts 2D landmark annotations to 3D and unifies all existing datasets, leading to the creation of LS3D-W, the largest and most challenging 3D facial landmark dataset to date ~230,000 images. (c) Following that, we train a neural network for 3D face alignment and evaluate it on the newly introduced LS3D-W. (d) We further look into the effect of all "traditional" factors affecting face alignment performance like large pose, initialization and resolution, and introduce a "new" one, namely the size of the network. (e) We show that both 2D and 3D face alignment networks achieve performance of remarkable accuracy which is probably close to saturating the datasets used. Training and testing code as well as the dataset can be downloaded from https://www.adrianbulat.com/face-alignment/
accepted to ICCV 2017
References in corpus (1)
Cited by in corpus (72)
- Res2Net: A New Multi-scale Backbone Architecture
- Artificial Intelligence in the Creative Industries: A Review
- A Comprehensive Analysis of Deep Regression
- MakeItTalk: Speaker-Aware Talking-Head Animation
- Learning Spatial Attention for Face Super-Resolution
- Visual Speech Recognition for Multiple Languages in the Wild
- On Learning 3D Face Morphable Model from In-the-wild Images
- PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
- SwinFace: A Multi-task Transformer for Face Recognition, Expression Recognition, Age Estimation and Attribute Estimation
- Neural Head Reenactment with Latent Pose Descriptors
- 3D Face Reconstruction from A Single Image Assisted by 2D Face Images in the Wild
- Multi-task head pose estimation in-the-wild
- Emotional Speech-Driven Animation with Content-Emotion Disentanglement
- Precise Facial Landmark Detection by Reference Heatmap Transformer
- Joint Face Image Restoration and Frontalization for Recognition
- SADRNet: Self-Aligned Dual Face Regression Networks for Robust 3D Dense Face Alignment and Reconstruction
- Self-supervised Learning of Detailed 3D Face Reconstruction
- Training Strategies for Improved Lip-reading
- Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
- Semi-Cycled Generative Adversarial Networks for Real-World Face Super-Resolution
- Facial-Sketch Synthesis: A New Challenge
- Supervision by Registration and Triangulation for Landmark Detection
- The Multimodal Driver Monitoring Database: A Naturalistic Corpus to Study Driver Attention
- Heatmap Regression via Randomized Rounding
- Face Hallucination with Finishing Touches
- Towards 3D Face Reconstruction in Perspective Projection: Estimating 6DoF Face Pose from Monocular Image
- Hierarchical binary CNNs for landmark localization with limited resources
- Talking Face Generation by Adversarially Disentangled Audio-Visual Representation
- Multi-view consensus CNN for 3D facial landmark placement
- Prior-Guided Multi-View 3D Head Reconstruction
- FLARE: Fast Learning of Animatable and Relightable Mesh Avatars
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
- A Full Transformer-based Framework for Automatic Pain Estimation using Videos
- Pro-UIGAN: Progressive Face Hallucination from Occluded Thumbnails
- Deep Face Video Inpainting via UV Mapping
- FreeStyleGAN: Free-view Editable Portrait Rendering with the Camera Manifold
- HeadPosr: End-to-end Trainable Head Pose Estimation using Transformer Encoders
- A Deep Learning Framework to Reconstruct Face under Mask
- SARGAN: Spatial Attention-based Residuals for Facial Expression Manipulation
- IFQA: Interpretable Face Quality Assessment
- Lite2Relight: 3D-aware Single Image Portrait Relighting
- From the Lab to the Wild: Affect Modeling via Privileged Information
- VR Facial Animation for Immersive Telepresence Avatars
- Multi-NeuS: 3D Head Portraits from Single Image with Neural Implicit Functions
- Automated Speaker Independent Visual Speech Recognition: A Comprehensive Survey
- A vector quantized masked autoencoder for audiovisual speech emotion recognition
- Animating Through Warping: an Efficient Method for High-Quality Facial Expression Animation
- Instance-level Facial Attributes Transfer with Geometry-Aware Flow
- Do Inpainting Yourself: Generative Facial Inpainting Guided by Exemplars
- Versatile Face Animator: Driving Arbitrary 3D Facial Avatar in RGBD Space
- Infinite 3D Landmarks: Improving Continuous 2D Facial Landmark Detection
- Single-Camera 3D Head Fitting for Mixed Reality Clinical Applications
- An Effective Deep Network for Head Pose Estimation without Keypoints
- Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis
- Evaluating User Experience and Data Quality in Gamified Data Collection for Appearance-Based Gaze Estimation
- Speaker-Adapted End-to-End Visual Speech Recognition for Continuous Spanish
- Interpretable and Flexible Target-Conditioned Neural Planners For Autonomous Vehicles
- Addressing Data Scarcity in Multimodal User State Recognition by Combining Semi-Supervised and Supervised Learning
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual Priors
- SOAP: Style-Omniscient Animatable Portraits
- Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
- FacialFilmroll: High-resolution multi-shot video editing
- Learning 3D Face Reconstruction with a Pose Guidance Network
- Impact of facial landmark localization on facial expression recognition
- Pixel Sampling for Style Preserving Face Pose Editing
- Teacher-Student Network for Real-World Face Super-Resolution with Progressive Embedding of Edge Information
- A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
- High-Quality Facial Albedo Generation for 3D Face Reconstruction from a Single Image using a Coarse-to-Fine Approach
- Facial 3D Model Registration Under Occlusions With SensiblePoints-based Reinforced Hypothesis Refinement
- HOT-POT: Optimal Transport for Sparse Stereo Matching
- Data-driven Head Motion Generation through Natural Gaze-Head Coordination
- Improving Generative Adversarial Network Generalization for Facial Expression Synthesis