1 paper
Siheng Xiong, Joe Zou, Faramarz Fekri +1
The quadratic cost of attention limits the scalability of long-context LLMs, especially under limited hardware memory budgets. While attention is often sparse, existing static spar…