1 paper · 1 filter
Zhanchao Xu, Haoyang Li, Qingfa Xiao +4
Existing sparse attention and KV cache compression methods for long-context LLM inference typically apply fixed sparsity patterns or uniform budgets across all attention heads, ove…