1 paper
Soosung Kim, Minjae Park, Eui-Young Chung +1
The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory. Vector Quantization (V…