1 paper · 1 filter
Siheng Xiong, Joe Zou, Faramarz Fekri +1
The quadratic cost of attention limits the scalability of long-context LLMs, especially under limited hardware memory budgets. While attention is often sparse, existing static spar…