1 paper · 1 filter
Dong Liu, Yanxuan Yu
Serving large language models (LLMs) efficiently remains challenging due to the high memory and latency overhead of key-value (KV) cache access during autoregressive decoding. We p…