Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks
arXiv:1604.02878 · doi:10.1109/LSP.2016.2603342
Abstract
Face detection and alignment in unconstrained environment are challenging due to various poses, illuminations and occlusions. Recent studies show that deep learning approaches can achieve impressive performance on these two tasks. In this paper, we propose a deep cascaded multi-task framework which exploits the inherent correlation between them to boost up their performance. In particular, our framework adopts a cascaded structure with three stages of carefully designed deep convolutional networks that predict face and landmark location in a coarse-to-fine manner. In addition, in the learning process, we propose a new online hard sample mining strategy that can improve the performance automatically without manual sample selection. Our method achieves superior accuracy over the state-of-the-art techniques on the challenging FDDB and WIDER FACE benchmark for face detection, and AFLW benchmark for face alignment, while keeps real time performance.
Submitted to IEEE Signal Processing Letters
Cited by in corpus (175)
- A Survey of the Recent Architectures of Deep Convolutional Neural Networks
- Deep Facial Expression Recognition: A Survey
- Additive Margin Softmax for Face Verification
- Facial Expression Recognition with Visual Transformers and Attentional Selective Fusion
- Combining EfficientNet and Vision Transformers for Video Deepfake Detection
- Biometric Face Presentation Attack Detection with Multi-Channel Convolutional Neural Network
- Deepfake Detection by Human Crowds, Machines, and Machine-informed Crowds
- Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition
- On the Reconstruction of Face Images from Deep Face Templates
- Learning Spatial Attention for Face Super-Resolution
- EVM-CNN: Real-Time Contactless Heart Rate Estimation from Facial Video
- Vision Transformer with Attentive Pooling for Robust Facial Expression Recognition
- AutoHR: A Strong End-to-end Baseline for Remote Heart Rate Measurement with Neural Searching
- 6D Rotation Representation For Unconstrained Head Pose Estimation
- FoodNet: Recognizing Foods Using Ensemble of Deep Networks
- SwinFace: A Multi-task Transformer for Face Recognition, Expression Recognition, Age Estimation and Attribute Estimation
- End2End Occluded Face Recognition by Masking Corrupted Features
- Multispectral Video Fusion for Non-contact Monitoring of Respiratory Rate and Apnea
- Learning Meta Pattern for Face Anti-Spoofing
- Evolving Boxes for Fast Vehicle Detection
- Efficient Facial Representations for Age, Gender and Identity Recognition in Organizing Photo Albums using Multi-output CNN
- Stimuli-Aware Visual Emotion Analysis
- One Detector to Rule Them All: Towards a General Deepfake Attack Detection Framework
- FSNet: An Identity-Aware Generative Model for Image-based Face Swapping
- Deep Learning Framework to Detect Face Masks from Video Footage
- Towards End-to-End Face Recognition through Alignment Learning
- When Age-Invariant Face Recognition Meets Face Age Synthesis: A Multi-Task Learning Framework and A New Benchmark
- PFA-GAN: Progressive Face Aging with Generative Adversarial Network
- Robust Subspace Clustering with Compressed Data
- HEU Emotion: A Large-scale Database for Multi-modal Emotion Recognition in the Wild
- MIDV-2020: A Comprehensive Benchmark Dataset for Identity Document Analysis
- Cross-Quality LFW: A Database for Analyzing Cross-Resolution Image Face Recognition in Unconstrained Environments
- Disentangling Identity and Pose for Facial Expression Recognition
- PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship
- Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
- Detect Faces Efficiently: A Survey and Evaluations
- Facial-Sketch Synthesis: A New Challenge
- Head2Head++: Deep Facial Attributes Re-Targeting
- Minimum Margin Loss for Deep Face Recognition
- CoReD: Generalizing Fake Media Detection with Continual Representation using Distillation
- EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion Embeddings
- Cross-Forgery Analysis of Vision Transformers and CNNs for Deepfake Image Detection
- Heatmap Regression via Randomized Rounding
- Pedestrian Detection with Wearable Cameras for the Blind: A Two-way Perspective
- Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model
- The MuSe 2022 Multimodal Sentiment Analysis Challenge: Humor, Emotional Reactions, and Stress
- An End-to-End Review of Gaze Estimation and its Interactive Applications on Handheld Mobile Devices
- One Label, One Billion Faces: Usage and Consistency of Racial Categories in Computer Vision
- Convolutional Ordinal Regression Forest for Image Ordinal Estimation
- Identity and Posture Recognition in Smart Beds with Deep Multitask Learning
- LatentAvatar: Learning Latent Expression Code for Expressive Neural Head Avatar
- Noisy Student Training using Body Language Dataset Improves Facial Expression Recognition
- On adversarial patches: real-world attack on ArcFace-100 face recognition system
- Estimation of BMI from Facial Images using Semantic Segmentation based Region-Aware Pooling
- Biometric Backdoors: A Poisoning Attack Against Unsupervised Template Updating
- Anchor Cascade for Efficient Face Detection
- Polarimetric Thermal to Visible Face Verification via Attribute Preserved Synthesis
- Multi-Scale Thermal to Visible Face Verification via Attribute Guided Synthesis
- Who, Where, and What to Wear? Extracting Fashion Knowledge from Social Media
- Polarimetric Thermal to Visible Face Verification via Self-Attention Guided Synthesis
- LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading
- Re-identification of Individuals in Genomic Datasets Using Public Face Images
- Unconstrained Face Detection and Open-Set Face Recognition Challenge
- A Survey on Masked Facial Detection Methods and Datasets for Fighting Against COVID-19
- Deep Appearance Models: A Deep Boltzmann Machine Approach for Face Modeling
- Zooming Into Video Conferencing Privacy and Security Threats
- Maximum A Posteriori Estimation of Distances Between Deep Features in Still-to-Video Face Recognition
- AGRNet: Adaptive Graph Representation Learning and Reasoning for Face Parsing
- A Master Key Backdoor for Universal Impersonation Attack against DNN-based Face Verification
- Classification of Cognitive Load and Expertise for Adaptive Simulation using Deep Multitask Learning
- Versatile audio-visual learning for emotion recognition
- Reconfigurable Cyber-Physical System for Critical Infrastructure Protection in Smart Cities via Smart Video-Surveillance
- On Generating Identifiable Virtual Faces
- Unconstrained Face-Mask & Face-Hand Datasets: Building a Computer Vision System to Help Prevent the Transmission of COVID-19
- M-VAD Names: a Dataset for Video Captioning with Naming
- Look globally, age locally: Face aging with an attention mechanism
- GP-GAN: Gender Preserving GAN for Synthesizing Faces from Landmarks
- Incremental Deep Learning for Robust Object Detection in Unknown Cluttered Environments
- EvilModel 2.0: Bringing Neural Network Models into Malware Attacks
- Families In Wild Multimedia: A Multimodal Database for Recognizing Kinship
- SFA: Small Faces Attention Face Detector
- A Full Transformer-based Framework for Automatic Pain Estimation using Videos
- Accelerating Approximate Aggregation Queries with Expensive Predicates
- Pro-UIGAN: Progressive Face Hallucination from Occluded Thumbnails
- Uncertain Label Correction via Auxiliary Action Unit Graphs for Facial Expression Recognition
- Hybrid Multimodal Fusion for Humor Detection
- DFGC 2021: A DeepFake Game Competition
- Leveraging Diffusion For Strong and High Quality Face Morphing Attacks
- Facial Synthesis from Visual Attributes via Sketch using Multi-Scale Generators
- On the representation and methodology for wide and short range head pose estimation
- A Dataset and Benchmark Towards Multi-Modal Face Anti-Spoofing Under Surveillance Scenarios
- Face Mask Extraction in Video Sequence
- Black-Box Face Recovery from Identity Features
- Controllable Generation with Text-to-Image Diffusion Models: A Survey
- Second Edition FRCSyn Challenge at CVPR 2024: Face Recognition Challenge in the Era of Synthetic Data
- Second FRCSyn-onGoing: Winning Solutions and Post-Challenge Analysis to Improve Face Recognition with Synthetic Data
- MetaAge: Meta-Learning Personalized Age Estimators
- Face.evoLVe: A High-Performance Face Recognition Library
- Real-time AdaBoost cascade face tracker based on likelihood map and optical flow
- Facial UV Map Completion for Pose-invariant Face Recognition: A Novel Adversarial Approach based on Coupled Attention Residual UNets
- HeadPosr: End-to-end Trainable Head Pose Estimation using Transformer Encoders
- Open-set Face Recognition for Small Galleries Using Siamese Networks
- Heterogeneous Face Frontalization via Domain Agnostic Learning
- NimbRo-OP2X: Affordable Adult-sized 3D-printed Open-Source Humanoid Robot for Research
- Twins-PainViT: Towards a Modality-Agnostic Vision Transformer Framework for Multimodal Automatic Pain Assessment using Facial Videos and fNIRS
- Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
- AdvFAS: A robust face anti-spoofing framework against adversarial examples
- Kinship Verification Based on Cross-Generation Feature Interaction Learning
- Octuplet Loss: Make Face Recognition Robust to Image Resolution
- Towards a General Deep Feature Extractor for Facial Expression Recognition
- Volumetric landmark detection with a multi-scale shift equivariant neural network
- Synthetic Thermal and RGB Videos for Automatic Pain Assessment utilizing a Vision-MLP Architecture
- Attacking Face Recognition with T-shirts: Database, Vulnerability Assessment and Detection
- Sentiment Analysis of Fashion Related Posts in Social Media
- Automatic Group Cohesiveness Detection With Multi-modal Features
- Real-Time Shape Tracking of Facial Landmarks
- Impact of Facial Tattoos and Paintings on Face Recognition Systems
- Privacy-sensitive Objects Pixelation for Live Video Streaming
- Social Media Authentication and Combating Deepfakes using Semi-fragile Invisible Image Watermarking
- Effectiveness of State-of-the-Art Super Resolution Algorithms in Surveillance Environment
- DeePLT: Personalized Lighting Facilitates by Trajectory Prediction of Recognized Residents in the Smart Home
- Ulixes: Facial Recognition Privacy with Adversarial Machine Learning
- Deep Collaborative Multi-Modal Learning for Unsupervised Kinship Estimation
- Universal Adversarial Spoofing Attacks against Face Recognition
- Pairwise Emotional Relationship Recognition in Drama Videos: Dataset and Benchmark
- A Survey on Personalized Content Synthesis with Diffusion Models
- Work-Efficient Parallel Non-Maximum Suppression Kernels
- Joint Feature Distribution Alignment Learning for NIR-VIS and VIS-VIS Face Recognition
- Transform, Contrast and Tell: Coherent Entity-Aware Multi-Image Captioning
- A Multi-task Adversarial Attack Against Face Authentication
- Autonomous Learning for Face Recognition in the Wild via Ambient Wireless Cues
- Headset: Human emotion awareness under partial occlusions multimodal dataset
- Arbitrary-Resolution and Arbitrary-Scale Face Super-Resolution with Implicit Representation Networks
- Accessing Passersby Proxemic Signals through a Head-Worn Camera: Opportunities and Limitations for the Blind
- Unified Face Matching and Physical-Digital Spoofing Attack Detection
- A Driver Fatigue Recognition Algorithm Based on Spatio-Temporal Feature Sequence
- Prototype Memory for Large-scale Face Representation Learning
- Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
- Infinite 3D Landmarks: Improving Continuous 2D Facial Landmark Detection
- A Comparative Study of Calibration Methods for Imbalanced Class Incremental Learning
- A Cluster-Matching-Based Method for Video Face Recognition
- Neonatal Face and Facial Landmark Detection from Video Recordings
- Can deepfakes be created by novice users?
- CAST: Cross-Attentive Spatio-Temporal feature fusion for deepfake detection
- Development of a face mask detection pipeline for mask-wearing monitoring in the era of the COVID-19 pandemic: A modular approach
- FaceTouch: Detecting hand-to-face touch with supervised contrastive learning to assist in tracing infectious disease
- SynMorph: Generating Synthetic Face Morphing Dataset with Mated Samples
- ID-Booth: Identity-consistent Face Generation with Diffusion Models
- Jet Single Shot Detection
- Bounding-box deep calibration for high performance face detection
- How Unique Is a Face: An Investigative Study
- An Efficient Method for Face Quality Assessment on the Edge
- Video Sentiment Analysis with Bimodal Information-augmented Multi-Head Attention
- End-to-end Evaluation of Practical Video Analytics Systems for Face Detection and Recognition
- User profiles matching for different social networks based on faces embeddings
- Expression-aware video inpainting for HMD removal in XR applications
- AI as a Tool for Fair Journalism: Case Studies from Malta
- The Face of Populism: Examining Differences in Facial Emotional Expressions of Political Leaders Using Machine Learning
- From Pixels to Words: Leveraging Explainability in Face Recognition through Interactive Natural Language Processing
- Practical X-ray Gastric Cancer Diagnostic Support Using Refined Stochastic Data Augmentation and Hard Boundary Box Training
- Continuous Facial Motion Deblurring
- Improving Facial Emotion Recognition through Dataset Merging and Balanced Training Strategies
- Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
- A Masked Face Classification Benchmark on Low-Resolution Surveillance Images
- Open-Set Face Identification on Few-Shot Gallery by Fine-Tuning
- NaMemo: Enhancing Lecturers' Interpersonal Competence of Remembering Students' Names
- FaceTracer: Unveiling Source Identities from Swapped Face Images and Videos for Fraud Prevention
- Generating 2D and 3D Master Faces for Dictionary Attacks with a Network-Assisted Latent Space Evolution
- Set-Based Face Recognition Beyond Disentanglement: Burstiness Suppression With Variance Vocabulary
- Learning Continuous Face Representation with Explicit Functions
- A Clustering-Based Method for Automatic Educational Video Recommendation Using Deep Face-Features of Lecturers
- Can we truly transfer an actor's genuine happiness to avatars? An investigation into virtual, real, posed and spontaneous faces
- Phase-shifted remote photoplethysmography for estimating heart rate and blood pressure from facial video
- Domain-generalizable Face Anti-Spoofing with Patch-based Multi-tasking and Artifact Pattern Conversion
- An Empirical Study of Visual Features for DNN based Audio-Visual Speech Enhancement in Multi-talker Environments