1 paper
Puyuan Peng, Shang-Wen Li, Abdelrahman Mohamed +1
We present VoiceStar, the first zero-shot TTS model that achieves both output duration control and extrapolation. VoiceStar is an autoregressive encoder-decoder neural codec langua…