1 paper · 1 filter
Yuqi Pan, Zheng Li, Bohao Tang +2
As context lengths scale, attention increasingly becomes a primary computational bottleneck in large language models. Standard Transformers remain powerful but computationally inef…