1 paper
Ambuj Mehrish, Abhinaba Roy, Alex Ivanov +2
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We i…