1 paper
Shun Lei, Yixuan Zhou, Liyang Chen +8
Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker's voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and mod…