1 paper
Matthew Baas, Pieter Scholtz, Arnav Mehta +3
Codec-based text-to-speech (TTS) models have shown impressive quality with zero-shot voice cloning abilities. However, they often struggle with more expressive references or comple…