1 paper
Jinhan Chen, Jianchun Liu, Hongli Xu +2
The growing memory footprint of the Key-Value (KV) cache poses a severe scalability bottleneck for long-context Large Language Model (LLM) inference. While KV cache eviction has em…