1 paper · 1 filter
Shi Qiu, Yifan Hu, Xintao Wang +6
LLM serving relies on prefix caching to improve inference performance. As growing contexts push key-value (KV) cache footprint far beyond GPU HBM and CPU DRAM capacity, KV cache is…