3 papers
cs.SD2025
EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
Jiacheng Shi, Hongfei Du, Y. Alicia Hong +1
Speech emotion recognition (SER) with audio-language models (ALMs) remains vulnerable to distribution shifts at test time, leading to performance degradation in out-of-domain scena…
cs.AI2025
Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
Jiacheng Shi, Hongfei Du, Y. Alicia Hong +1
Large audio-language models (LALMs) exhibit strong zero-shot performance across speech tasks but struggle with speech emotion recognition (SER) due to weak paralinguistic modeling…
cs.CL2025
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
Jiacheng Shi, Hongfei Du, Yangfan He +2
Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utter…