1 paper · 1 filter
Peiyu Liu, Ze-Feng Gao, Wayne Xin Zhao +3
Key-value~(KV) caching is an important technique to accelerate the inference of large language models~(LLMs), but incurs significant memory overhead. To compress the size of KV cac…