Causal Supervision of Attention for Affective Behaviour Analysis
arXiv:2607.12091
The paper introduces a causal supervision and cross‑covariance regularization scheme for attention pooling that encourages subject‑invariant facial attention, improving multi‑task affective behavior analysis (valence‑arousal, expression, and action unit detection).
Abstract
The \textit{11th Affective Behaviour Analysis in-the-wild Competition} includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves for VA estimation on the official validation set, together with and for expression recognition and action unit detection, respectively, resulting in an overall score (the sum of the individual task metrics) of .
10 pages, 1 figure, 2 tables