1 paper · 1 filter
Dongjie Xu, Kai Qian, Julius +6
Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as l…