1 paper · 1 filter
Libo Sun, Peixiong He, Po-Wei Harn +1
KV cache memory is the dominant bottleneck for long-context LLM inference. Existing compression methods each act on a single axis of the four-dimensional KV tensor -- token evictio…