computer vision

A Calibrated Multimodal Ensemble for Ambivalence/Hesitancy Recognition: System Description and Private-Test Submission Strategy

arXiv:2607.12176

summary

The paper presents a calibrated, equal‑weight ensemble of three multimodal fusion models that combine frozen face, audio, text, and pose embeddings to automatically recognize ambivalence and hesitancy in video, achieving strong macro‑F1 scores on public and private test sets.

Abstract

Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the ABAW A/H challenge on the BAH dataset. We describe our system for the 11th edition of the challenge: a calibrated, equal-weight ensemble of three fusion models over frozen face, audio, text, and pose embeddings, which reaches 0.7358 macro-F1 on the public test set. This year's private test, released on a disjoint set of 30 new participants, is scored on five allowed submissions; we report the configuration and rationale of each of our five submissions, and, where already available, the private-test score obtained. Our first submission, an exact replica of the calibrated ensemble tuned only on public validation, scored 0.7361 macro-F1 on the private test, matching our public-test estimate almost exactly and confirming the pipeline generalizes to unseen participants without leakage.

8 pages, 1 figure, ECCV workshops

Topics & keywords

#multimodal fusion#ambivalence detection#hesitancy recognition#video affect analysis#ensemble learningface embeddingsaudio embeddingstext embeddingspose embeddingscalibrated ensemblemacro-F1
A Calibrated Multimodal Ensemble for Ambivalence/Hesitancy Recognition: System Description and Private-Test Submission Strategy · wovepaper