1 paper
Fengfeng Liang, Yuechen Zhang, Jiaya Jia
Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contribution to a future attention logit decomposes into a position-…