1 paper
Maanas Taneja, Purab Shingvi
The key-value (KV) cache in large language models presents a significant memory bottleneck during inference, growing linearly with sequence length and often exceeding the memory fo…