1 paper · 1 filter
Junjie Li, Jiong Lou, Jie Li
Efficient long-context inference in Large Language Models (LLMs) is severely constrained by the Key-Value (KV) cache memory wall, yet existing pruning methods force a choice betwee…