1 paper
Chunan Shi, Yilei Chen, Yilin Chen +2
Large Language Model (LLM) inference relies on key-value (KV) caches to avoid redundant attention computation. While approximate KV cache retention techniques reduce memory usage b…