1 paper
Juncheng Wang, Zhe Hu, Chao Xu +5
Autoregressive (AR) models excel at generating temporally coherent audio by producing tokens sequentially, yet they often falter in faithfully following complex textual prompts, es…