Do Deep Neural Networks Learn Facial Action Units When Doing Expression Recognition?
arXiv:1510.02969
Abstract
Despite being the appearance-based classifier of choice in recent years, relatively few works have examined how much convolutional neural networks (CNNs) can improve performance on accepted expression recognition benchmarks and, more importantly, examine what it is they actually learn. In this work, not only do we show that CNNs can achieve strong performance, but we also introduce an approach to decipher which portions of the face influence the CNN's predictions. First, we train a zero-bias CNN on facial expression data and achieve, to our knowledge, state-of-the-art performance on two expression recognition benchmarks: the extended Cohn-Kanade (CK+) dataset and the Toronto Face Dataset (TFD). We then qualitatively analyze the network by visualizing the spatial patterns that maximally excite different neurons in the convolutional layers and show how they resemble Facial Action Units (FAUs). Finally, we use the FAU labels provided in the CK+ dataset to verify that the FAUs observed in our filter visualizations indeed align with the subject's facial movements.
Accepted at ICCV 2015 CV4AC Workshop. Corrected numbers in Tables 2 and 3
Cited by in corpus (12)
- Deep Facial Expression Recognition: A Survey
- Facial Expression Recognition using Convolutional Neural Networks: State of the Art
- A Deep Learning Perspective on the Origin of Facial Expressions
- Deep generative-contrastive networks for facial expression recognition
- CAKE: Compact and Accurate K-dimensional representation of Emotion
- From Facial Expression Recognition to Interpersonal Relation Prediction
- How Deep Neural Networks Can Improve Emotion Recognition on Video Data
- A Scalable Approach for Facial Action Unit Classifier Training UsingNoisy Data for Pre-Training
- Robust Emotion Recognition from Low Quality and Low Bit Rate Video: A Deep Learning Approach
- Deep Multi-Facial Patches Aggregation Network For Facial Expression Recognition
- DeepCoder: Semi-parametric Variational Autoencoders for Automatic Facial Action Coding
- Efficient Facial Feature Learning with Wide Ensemble-based Convolutional Neural Networks