1 paper · 1 filter
Heng Wang, Jielin Qiu, Wenting Zhao +7
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV ca…