1 paper · 1 filter
Junzhe Yang, Xiaoyu Shen
The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate this by evicting tokens based on importa…