6 papers · 1 filter
Beyond Reconstruction: Full-Context Generative DiT for Music Generation
Yunjia Li, Menglin Wu, Junyu Dai +13
Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are tra…
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scene…
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
Bajian Xiang, Cheng Wen, Han Zhao +12
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, au…
Audio Deep Fake Detection System with Neural Stitching for ADD 2022
Rui Yan, Cheng Wen, Shuran Zhou +3
This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two ro…
Time Domain Adversarial Voice Conversion for ADD 2022
Cheng Wen, Tingwei Guo, Xingjun Tan +5
In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) sy…
DiDiSpeech: A Large Scale Mandarin Speech Corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang +8
This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the…