1 paper
Yuhang Zhan, Lisi Chen, Shuo Shang
KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate cont…