7 papers · 1 filter
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Xiutian Zhao, Luqi Sun, Björn Schuller +1
Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. How…
Multilingual Emotion Neurons in Large Audio-Language Models
Xiutian Zhao, Philipp Koehn, Björn Schuller +1
Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks,…
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
Yupei Li, Qiyang Sun, Xiaoliang Wu +3
Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explai…
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn +2
Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic…
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
Xiutian Zhao, Björn Schuller, Björn Schuller +1
Emotion is a central dimension of spoken communication, yet, we still lack a mechanistic account of how modern large audio-language models (LALMs) encode it internally. We present…
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
Zhenqi Jia, Rui Liu, Berrak Sisman +1
Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate pro…