1 paper
Namgyu Ho, Huzama Ahmad, Woosung Koh +3
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previou…