1 paper
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi +1
We demonstrate that unstructured sparsity significantly improves KV cache compression for LLMs, enabling sparsity levels up to 70% without compromising accuracy or requiring fine-t…