collaborators

14 papers

cs.SD2025

Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought

Zhixian Zhao, Xinfa Zhu, Xinsheng Wang +4

Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. Howeve…

eess.AS2025

DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching

Hanke Xie, Dake Guo, Chengyou Wang +8

Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However,…

cs.CL2025

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…

eess.AS2025

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

Kangxiang Xia, Xinfa Zhu, Jixun Yao +1

In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has pro…

eess.AS2025

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation

Tianlun Zuo, Jingbin Hu, Yuke Li +6

Zero-shot emotion transfer in cross-lingual speech synthesis refers to generating speech in a target language, where the emotion is expressed based on reference speech from a diffe…

cs.MM2025

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

Wenjie Tian, Xinfa Zhu, Haohe Liu +6

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This…