1 paper
Weiye Wang, Chen Chen, Junxue Zhang +7
Distributed prefix caching has become a core technique for efficient LLM serving. However, for long-context requests with high cache hit ratios, retrieving reusable KVCache blocks…