1 paper
Bowen Peng, Subho Ghosh, Jeffrey Quesnelle
Training causal transformers at extreme sequence lengths is bottlenecked by the quadratic time and memory of scaled dot-product attention (SDPA). In this work, we propose Lighthous…