47 citations · 149 across the 15 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2024★ 1 cited
RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging
Mingyang Zhang, Yi Zhou, Yi Ren +3
This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally t…
eess.AS2023★ 16 cited
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Ziyue Jiang, Yi Ren, Zhenhui Ye +9
Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. Howeve…
eess.AS2022★ 21 cited
ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
Rongjie Huang, Zhou Zhao, Huadai Liu +3
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinde…