1 paper
Daeha Lee, Do-Hyung Kim, Jae-Hong Kim
The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quant…