1 paper · 1 filter
Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on…