1 paper
Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1
Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to f…