Facial Action Unit Detection via Adaptive Attention and Relation
arXiv:2001.01168 · doi:10.1109/TIP.2023.3277794
Abstract
Facial action unit (AU) detection is challenging due to the difficulty in capturing correlated information from subtle and dynamic AUs. Existing methods often resort to the localization of correlated regions of AUs, in which predefining local AU attentions by correlated facial landmarks often discards essential parts, or learning global attention maps often contains irrelevant areas. Furthermore, existing relational reasoning methods often employ common patterns for all AUs while ignoring the specific way of each AU. To tackle these limitations, we propose a novel adaptive attention and relation (AAR) framework for facial AU detection. Specifically, we propose an adaptive attention regression network to regress the global attention map of each AU under the constraint of attention predefinition and the guidance of AU detection, which is beneficial for capturing both specified dependencies by landmarks in strongly correlated regions and facial globally distributed dependencies in weakly correlated regions. Moreover, considering the diversity and dynamics of AUs, we propose an adaptive spatio-temporal graph convolutional network to simultaneously reason the independent pattern of each AU, the inter-dependencies among AUs, as well as the temporal dependencies. Extensive experiments show that our approach (i) achieves competitive performance on challenging benchmarks including BP4D, DISFA, and GFT in constrained scenarios and Aff-Wild2 in unconstrained scenarios, and (ii) can precisely learn the regional correlation distribution of each AU.
This paper has been accepted by IEEE Transactions on Image Processing (TIP)
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment
- JA-Net: Joint Facial Action Unit Detection and Face Alignment via Adaptive Attention
- AU R-CNN: Encoding Expert Prior Knowledge into R-CNN for Action Unit Detection
- Facial Action Unit Detection Using Attention and Relation Learning
- Relation Modeling with Graph Convolutional Networks for Facial Action Unit Detection
- GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
Cited by in corpus (4)
- Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
- Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
- Micro-Expression Recognition via Fine-Grained Dynamic Perception
- Symmetric Perception and Ordinal Regression for Detecting Scoliosis Natural Image