1 paper · 1 filter
Michael Wang, Keith Li, Roozbeh Bostandoost
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. However, exist…