Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
Qwen3-TTS Technical Report
Hangrui Hu, Xinfa Zhu, Ting He +13
In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3…
cs.SD2025
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
Dingdong Wang, Jin Xu, Ruihang Chu +6
Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to s…
cs.SD2024
SongTrans: An unified song transcription and alignment method for lyrics and notes
Siwei Wu, Jinzheng He, Ruibin Yuan +5
The quantity of processed data is crucial for advancing the field of singing voice synthesis. While there are tools available for lyric or note transcription tasks, they all need p…