1 paper
Boyu Feng, Jiahong Liu, Yifan Li +7
Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and deco…