computer vision

Emotion Recognition in Signers

arXiv:2512.15376

summary

The paper introduces a new benchmark for recognizing emotions in Japanese Sign Language signers and shows how using spoken-language text data, careful temporal segment selection, and hand‑motion information can improve emotion classification despite limited sign language data.

Abstract

Recognition of signers' emotions suffers from one theoretical challenge and one practical challenge, namely, the overlap between grammatical and affective facial expressions and the scarcity of data for model training. This paper addresses these two challenges in a cross-lingual setting using our eJSL dataset, a new benchmark dataset for emotion recognition in Japanese Sign Language signers, and BOBSL, a large British Sign Language dataset with subtitles. In eJSL, two signers expressed 78 distinct utterances with each of seven different emotional states, resulting in 1,092 video clips. We empirically demonstrate that 1) textual emotion recognition in spoken language mitigates data scarcity in sign language, 2) temporal segment selection has a significant impact, and 3) incorporating hand motion enhances emotion recognition in signers. Finally we establish a stronger baseline than spoken language LLMs.

Topics & keywords

#sign language#emotion recognition#cross-lingual learning#temporal segmentation#hand motioneJSL datasetBOBSL datasettextual emotion transferhand motion featureslarge language model baseline
Emotion Recognition in Signers · wovepaper