1 paper
Qiankun Ma, Yanjiang Zhou, Zinan Xiong +5
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-re…