1 paper
Rahul Krishnan, Volker Schulz
The key-value (KV) cache has become the dominant memory cost of transformer inference: it grows with batch size, context length, and depth, and at long context it, rather than the…