1 paper · 1 filter
Zunhai Su, Zhe Chen, Wang Shen +4
Key-Value (KV) cache facilitates efficient large language models (LLMs) inference by avoiding recomputation of past KVs. As the batch size and context length increase, the oversize…