1 paper
Ke Lin, Yiyang Luo, Zhaolong Su +2
Causal Transformer language models suffer from strictly sequential decoding and a quadratic per-step attention cost. While linear-time causal models and discrete diffusion models e…