Robust Face Recognition via Multimodal Deep Face Representation
arXiv:1509.00244 · doi:10.1109/TMM.2015.2477042
Abstract
Face images appeared in multimedia applications, e.g., social networks and digital entertainment, usually exhibit dramatic pose, illumination, and expression variations, resulting in considerable performance degradation for traditional face recognition algorithms. This paper proposes a comprehensive deep learning framework to jointly learn face representation using multimodal information. The proposed deep learning structure is composed of a set of elaborately designed convolutional neural networks (CNNs) and a three-layer stacked auto-encoder (SAE). The set of CNNs extracts complementary facial features from multimodal data. Then, the extracted features are concatenated to form a high-dimensional feature vector, whose dimension is compressed by SAE. All the CNNs are trained using a subset of 9,000 subjects from the publicly available CASIA-WebFace database, which ensures the reproducibility of this work. Using the proposed single CNN architecture and limited training data, 98.43% verification rate is achieved on the LFW database. Benefited from the complementary information contained in multimodal data, our small ensemble system achieves higher than 99.0% recognition rate on LFW using publicly available training set.
To appear in IEEE Trans. Multimedia
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Deep Learning Face Representation by Joint Identification-Verification
- Learning Face Representation from Scratch
- Classification with Noisy Labels by Importance Reweighting
- Multi-Directional Multi-Level Dual-Cross Patterns for Robust Face Recognition
- Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?
- Face Search at Scale: 80 Million Gallery
- A Comprehensive Survey on Pose-Invariant Face Recognition
- Web-Scale Training for Face Identification
Cited by in corpus (37)
- DehazeNet: An End-to-End System for Single Image Haze Removal
- Deep Face Recognition: A Survey
- Trunk-Branch Ensemble Convolutional Neural Networks for Video-based Face Recognition
- Large-Margin Softmax Loss for Convolutional Neural Networks
- SphereFace: Deep Hypersphere Embedding for Face Recognition
- Representation Learning by Rotating Your Faces
- Recent Advances in Deep Learning Techniques for Face Recognition
- Evaluation of an Audio-Video Multimodal Deepfake Dataset using Unimodal and Multimodal Detectors
- SiGAN: Siamese Generative Adversarial Network for Identity-Preserving Face Hallucination
- von Mises-Fisher Mixture Model-based Deep learning: Application to Face Verification
- Multi-modal Deep Analysis for Multimedia
- When Wireless Communications Meet Computer Vision in Beyond 5G
- A multimodal deep learning approach for named entity recognition from social media
- The Labeled Multiple Canonical Correlation Analysis for Information Fusion
- DeepIris: Iris Recognition Using A Deep Learning Approach
- FingerNet: Pushing The Limits of Fingerprint Recognition Using Convolutional Neural Network
- A Comprehensive Survey on Pose-Invariant Face Recognition
- Deep Appearance Models: A Deep Boltzmann Machine Approach for Face Modeling
- Multimodal Recurrent Neural Networks with Information Transfer Layers for Indoor Scene Labeling
- Survey on Deep Multi-modal Data Analytics: Collaboration, Rivalry and Fusion
- Face Recognition Using Deep Multi-Pose Representations
- Pasadena: Perceptually Aware and Stealthy Adversarial Denoise Attack
- Multi-Task Learning by Deep Collaboration and Application in Facial Landmark Detection
- Cascaded Asymmetric Local Pattern: A Novel Descriptor for Unconstrained Facial Image Recognition and Retrieval
- Ulixes: Facial Recognition Privacy with Adversarial Machine Learning
- A Detailed Look At CNN-based Approaches In Facial Landmark Detection
- RankPose: Learning Generalised Feature with Rank Supervision for Head Pose Estimation
- Adversarial Relighting Against Face Recognition
- Fully Learnable Group Convolution for Acceleration of Deep Neural Networks
- Facial Landmark Machines: A Backbone-Branches Architecture with Progressive Representation Learning
- DeMeshNet: Blind Face Inpainting for Deep MeshFace Verification
- An Interpretable Compression and Classification System: Theory and Applications
- PI-GAN: Learning Pose Independent representations for multiple profile face synthesis
- A Capsule-unified Framework of Deep Neural Networks for Graphical Programming
- Fast Landmark Localization with 3D Component Reconstruction and CNN for Cross-Pose Recognition
- ComplexFace: a Multi-Representation Approach for Image Classification with Small Dataset
- A Novel Perspective to Zero-shot Learning: Towards an Alignment of Manifold Structures via Semantic Feature Expansion