3 papers
cs.CL2026
COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
Weizhen Bian, Sitong Cheng, Rongxiu Zhong +9
Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on cl…
cs.SD2026
Qwen-Music Technical Report
Jin Xu, Kangdi Wang, Ruibin Yuan +24
We introduce Qwen-Music, a music generation model that produces high-fidelity songs with complete vocals. It supports text-to-music generation from descriptions, lyrics, and musica…
cs.SD2026
ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech
Wei Xue, Junlan Feng, Shilei Zhang +9
Recent advances in text-to-speech (TTS) have greatly improved speech naturalness, speaker similarity, and controllability. However, most existing controllable TTS systems still rel…