1 paper · 1 filter
Qiuyang Zhang, Kai Zhou, Ding Tang +5
Large language models encounter critical GPU memory capacity constraints during long-context inference, where KV cache memory consumption severely limits decode batch sizes. While…