1 paper
Jinzuomu Zhong, Korin Richmond, Zhiba Su +1
While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue,…