1 paper · 1 filter
Yifei Gao, Lei Wang, Rong-Cheng Tu +3
A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve…