1 paper
Renjie Xie, Juncheng Yang, Aoting Hu +4
Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains…