1 paper
Wei Gao, Xinyu Zhou, Peng Sun +2
Key-Value cache (\texttt{KV} \texttt{cache}) compression has emerged as a promising technique to optimize Large Language Model (LLM) serving. It primarily decreases the memory cons…