1 paper
Chao Fei, Guozhong Li, Chenxi Liu +1
Long-context LLMs demand accurate inference at low latency, yet decoding becomes primarily constrained by KV cache as context grows. Prior pruning methods are largely context-agnos…