3 papers
cs.SD2026
EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
Xin Jing, Andreas Triantafyllopoulos, Jiadong Wang +3
Recent advancements in speech captioning models have enabled the generation of rich, fine-grained captions for emotional speech. However, the evaluation of such captions remains a…
cs.SD2026
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
Xin Jing, Jiadong Wang, Andreas Triantafyllopoulos +4
The ambiguity of human emotions poses several challenges for machine learning models, as they often overlap and lack clear delineating boundaries. Contrastive language-audio pretra…
cs.AI2025
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
Xin Jing, Jiadong Wang, Iosif Tsangko +2
Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to…