8 papers
Multilingual Emotion Neurons in Large Audio-Language Models
Xiutian Zhao, Philipp Koehn, Björn Schuller +1
Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks,…
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion
Zihan Zhang, Shreeram Suresh Chandra, Zongyang Du +5
Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direc…
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews +2
To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing meth…
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn +2
Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic…
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
Weiting Tan, Andy T. Liu, Ming Tu +3
Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate…
Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
Weiting Tan, Xinghua Qu, Ming Tu +4
Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To t…