1 paper
Hengshuai Yao, Xing Chen, Ahmed Murtadha +8
Standard attention scales quadratically with sequence length. Efficient attention methods reduce this O(n^2) cost, but when retrofitted into pretrained models, they often degrade p…