1 paper
Heng Wang, Jielin Qiu, Wenting Zhao +7
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV ca…