1 paper
Gerard I. Gállego, Roy Fejgin, Chunghsin Yeh +2
Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplif…