1 paper
Zhen Huang, Ruizhe Yao, Danyi Liu +8
Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to acceler…