1 paper
Zhongwei Wan, Xinjian Wu, Yu Zhang +8
Generative inference in Large Language Models (LLMs) is impeded by the growing memory demands of Key-Value (KV) cache, especially for longer sequences. Traditional KV cache evictio…