1 paper · 1 filter
Zhiyuan Shi, Qibo Qiu, Feng Xue +5
The linear memory growth of the KV cache poses a significant bottleneck for LLM inference in long-context tasks. Existing static compression methods often fail to preserve globally…