1 paper · 1 filter
Hong Chen, Xiang Liu, Yubo Gao +5
Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank tokens by one importance si…