1 paper
Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1
The key-value (KV) cache is a major bottleneck in long-context inference, where memory and computation grow with sequence length. Existing KV eviction methods reduce this cost but…