1 paper
Shi Qiu, Yifan Hu, Xintao Wang +6
LLM serving relies on prefix caching to improve inference performance. As growing contexts push key-value (KV) cache footprint far beyond GPU HBM and CPU DRAM capacity, KV cache is…