activity
20202026
collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

Yunjia Li, Menglin Wu, Junyu Dai +13

Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are tra…

eess.AS2026

Qwen-Audio-3.0-Gen-Preview Technical Report

Junyu Dai, Xiaoyue Duan, Xinyue Fan +14

Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scene…

eess.AS2026

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

Bajian Xiang, Cheng Wen, Han Zhao +12

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, au…

eess.AS2022

Audio Deep Fake Detection System with Neural Stitching for ADD 2022

Rui Yan, Cheng Wen, Shuran Zhou +3

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two ro…

eess.AS2022

Time Domain Adversarial Voice Conversion for ADD 2022

Cheng Wen, Tingwei Guo, Xingjun Tan +5

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) sy…

eess.AS2020

DiDiSpeech: A Large Scale Mandarin Speech Corpus

Tingwei Guo, Cheng Wen, Dongwei Jiang +8

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the…