1 paper · 1 filter
Jack Shi, Jerry Gu
Every deployed sparse-attention or KV-cache-eviction rule keeps a subset of the keys, discards the rest, and renormalizes the attention weights over the kept set. Enumerating the e…