1 paper · 1 filter
Dongwei Wang, Zijie Liu, Song Wang +5
The Key-Value (KV) cache reading latency increases significantly with context lengths, hindering the efficiency of long-context LLM inference. To address this, previous works propo…