From the 1 of 4 linked papers with an AI index.
4 papers
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
Weihao Wu, Liang Cao, Xinyu Wu +4
Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create im…
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Weihao wu, Zhiwei Lin, Yixuan Zhou +6
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding o…
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
Jingbei Li, Sipan Li, Ping Chen +7
Automatic dubbing, which generates a corresponding version of the input speech in another language, could be widely utilized in many real-world scenarios such as video and game loc…