Trunk-Branch Ensemble Convolutional Neural Networks for Video-based Face Recognition
arXiv:1607.05427 · doi:10.1109/TPAMI.2017.2700390
Abstract
Human faces in surveillance videos often suffer from severe image blur, dramatic pose variations, and occlusion. In this paper, we propose a comprehensive framework based on Convolutional Neural Networks (CNN) to overcome challenges in video-based face recognition (VFR). First, to learn blur-robust face representations, we artificially blur training data composed of clear still images to account for a shortfall in real-world video training data. Using training data composed of both still images and artificially blurred data, CNN is encouraged to learn blur-insensitive features automatically. Second, to enhance robustness of CNN features to pose variations and occlusion, we propose a Trunk-Branch Ensemble CNN model (TBE-CNN), which extracts complementary information from holistic face images and patches cropped around facial components. TBE-CNN is an end-to-end model that extracts features efficiently by sharing the low- and middle-level convolutional layers between the trunk and branch networks. Third, to further promote the discriminative power of the representations learnt by TBE-CNN, we propose an improved triplet loss function. Systematic experiments justify the effectiveness of the proposed techniques. Most impressively, TBE-CNN achieves state-of-the-art performance on three popular video face databases: PaSC, COX Face, and YouTube Faces. With the proposed techniques, we also obtain the first place in the BTAS 2016 Video Person Recognition Evaluation.
Accepted Version to IEEE T-PAMI
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Deep Learning Face Representation by Joint Identification-Verification
- Learning Face Representation from Scratch
- DeepID3: Face Recognition with Very Deep Neural Networks
- Recent Progress in Image Deblurring
Cited by in corpus (32)
- Deep Face Recognition: A Survey
- Distributed Deep Learning Model for Intelligent Video Surveillance Systems with Edge Computing
- Hydra: an Ensemble of Convolutional Neural Networks for Geospatial Land Classification
- Examining the Impact of Blur on Recognition by Convolutional Networks
- Recent Advances in Deep Learning Techniques for Face Recognition
- Predicting Lung Nodule Malignancies by Combining Deep Convolutional Neural Network and Handcrafted Features
- Deep Learning Based Single Sample Per Person Face Recognition: A Survey
- Discriminative Adversarial Domain Generalization with Meta-learning based Cross-domain Validation
- A Reconfigurable Convolution-in-Pixel CMOS Image Sensor Architecture
- EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion Embeddings
- Embedding Visual Hierarchy with Deep Networks for Large-Scale Visual Recognition
- CDPM: Convolutional Deformable Part Models for Semantically Aligned Person Re-identification
- Maximum A Posteriori Estimation of Distances Between Deep Features in Still-to-Video Face Recognition
- Reasoning Graph Networks for Kinship Verification: from Star-shaped to Hierarchical
- FabricNet: A Fiber Recognition Architecture Using Ensemble ConvNets
- Asynchronous Federated Learning with Differential Privacy for Edge Intelligence
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Recurrent Embedding Aggregation Network for Video Face Recognition
- Video Face Recognition: Component-wise Feature Aggregation Network (C-FAN)
- An Automatic System for Unconstrained Video-Based Face Recognition
- Attention Control with Metric Learning Alignment for Image Set-based Recognition
- Modularity in Deep Learning: A Survey
- SeqFace: Make full use of sequence information for face recognition
- RRR-Net: Reusing, Reducing, and Recycling a Deep Backbone Network
- A Fast and Accurate System for Face Detection, Identification, and Verification
- Dual-Triplet Metric Learning for Unsupervised Domain Adaptation in Video-Based Face Recognition
- Finding Person Relations in Image Data of the Internet Archive
- SSDL: Self-Supervised Domain Learning for Improved Face Recognition
- Attention-Set based Metric Learning for Video Face Recognition
- A Joint Pixel and Feature Alignment Framework for Cross-dataset Palmprint Recognition
- Occlusion-guided compact template learning for ensemble deep network-based pose-invariant face recognition
- Feature Aggregation Network for Video Face Recognition