1 paper
Charles O'Neill, Alex Sandomirsky, Harry Partridge +2
The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call during inference, expressive…