1 paper
Fei Zuo, Zikang Zhou, Hao Cong +2
Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, making it a primary memory bottle…