1 paper
Ye Lu, Yihan Yan, Zhaoyang Zhang +4
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens s…