2 papers
eess.AS2024
M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
Shuai Wang, Pengcheng Zhu, Haizhou Li
Fixed-dimensional speaker embeddings have become the dominant approach in speaker modeling, typically spanning hundreds to thousands of dimensions. These dimensions are hyperparame…
eess.AS2024
E1 TTS: Simple and Fast Non-Autoregressive TTS
Zhijun Liu, Shuai Wang, Pengcheng Zhu +2
This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distributi…