14 papers
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
Zhixian Zhao, Xinfa Zhu, Xinsheng Wang +4
Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. Howeve…
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
Hanke Xie, Dake Guo, Chengyou Wang +8
Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However,…
Qwen3-Omni Technical Report
Jin Xu, Zhifang Guo, Hangrui Hu +35
We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
Kangxiang Xia, Xinfa Zhu, Jixun Yao +1
In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has pro…
XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
Tianlun Zuo, Jingbin Hu, Yuke Li +6
Zero-shot emotion transfer in cross-lingual speech synthesis refers to generating speech in a target language, where the emotion is expressed based on reference speech from a diffe…
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
Wenjie Tian, Xinfa Zhu, Haohe Liu +6
While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This…