1 paper · 1 filter
Gradwell Dzikanyanga, Yanqi Pan, Weihao Yang +3
Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads. Low-bit KV-cache quanti…