1 paper
Changfeng Gao, Yong Ren, Jun Yuan +3
Recent state-of-the-art (SOTA) text-to-speech (TTS) systems typically adopt a cascaded pipeline consisting of a speech tokenizer, an autoregressive large language model (LLM), and…