1 paper
Xu Yang, Jiapeng Zhang, Dongyang Zhao +2
The KV cache in self-attention has emerged as a major bottleneck in long-context and large-batch inference for LLMs. Existing approaches often treat sparsity prediction and compres…