1 paper
Yash Akhauri, Safeen Huda, Mohamed S. Abdelfattah
When predicting the next token in a sequence, vanilla transformers compute attention over all previous tokens, resulting in quadratic scaling of compute with sequence length. State…