4 papers · 1 filter
Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
Xiangyuan Xue, Jiajun Lu, Yan Gao +3
Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains cha…
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
Xiaofeng Yu, Jiaheng Dong, Jean Honorio +3
Speech emotion recognition plays an important role in various applications. However, most existing approaches predict a single emotion label, oversimplifying the inherently ambiguo…
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
Jule Valendo Halim, Siyi Wang, Hong Jia +1
Emotional intelligence in conversational AI is crucial across domains like human-computer interaction. While numerous models have been developed, they often overlook the complexity…
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
Ting Dang, Yan Gao, Hong Jia
Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processin…