1 paper
Ziyao Tang, Pengkun Jiao, Xinhang Chen +3
Given the quadratic complexity of attention, KV cache eviction is vital to accelerate model inference. Current KV cache eviction methods typically rely on instantaneous heuristic m…