1 paper
Songtao Liu, Hongwu Peng, Zhiwei Zhang +2
Long-context inference in large language models is bottlenecked by Key--Value (KV) cache loading during the decoding stage, where the sequential nature of generation requires repea…