1 paper
Yifan Zhang, Zhen Qin, Mengdi Wang +1
The quadratic cost of scaled dot-product attention is a central obstacle to scaling autoregressive language models to long contexts. Linear-time attention and State Space Models (S…