1 paper
Yuzhen Mao, Qitong Wang, Martin Ester +1
Key-Value (KV) cache plays a crucial role in accelerating inference in large language models (LLMs) by storing intermediate attention states and avoiding redundant computation duri…