4 papers
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
Bajian Xiang, Cheng Wen, Han Zhao +12
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, au…
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
Chengyuan Ma, Jiawei Jin, Ruijie Xiong +3
We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive au…
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
Jiawei Jin, Zhihan Yang, Yixuan Zhou +1
We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and envir…
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model
Lingzhou Mu, Baiji Liu, Ruonan Zhang +4
Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressio…