1 paper
Hong Chen, Yudong Zeng, Yongwei Huang +4
Long-context inference is bottlenecked by the memory footprint of the key-value (KV) cache, especially for small models under tight resource budgets. Existing KV cache eviction met…