4 papers
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
Yuepeng Jiang, Ziqian Ning, Shuai Wang +5
In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ens…
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
Sho Inoue, Shuai Wang, Wanxing Wang +3
In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, w…
M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
Shuai Wang, Pengcheng Zhu, Haizhou Li
Fixed-dimensional speaker embeddings have become the dominant approach in speaker modeling, typically spanning hundreds to thousands of dimensions. These dimensions are hyperparame…
E1 TTS: Simple and Fast Non-Autoregressive TTS
Zhijun Liu, Shuai Wang, Pengcheng Zhu +2
This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distributi…