1 paper
Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev +2
With the growing demand for long-context LLMs across a wide range of applications, the key-value (KV) cache has become a critical bottleneck for both latency and memory usage. Rece…