1 paper
Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen +4
Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allo…