1 paper
Sourjya Roy, Shrihari Sridharan, Surya Selvam +1
As Large Language Models (LLMs) scale in size and context length, the memory requirements of the key value (KV) cache have emerged as a major bottleneck during autoregressive decod…