1 paper
Jiawei Du, I-Ming Lin, I-Hsiang Chiu +6
Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, hu…