1 paper · 1 filter
Qiankun Ma, Yanjiang Zhou, Zinan Xiong +5
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-re…