1 paper · 1 filter
Xingyu Xiang, Raj Joshi, Yuhan Liu +6
Distributed prefix caching accelerates long-context LLM serving by reusing KV cache entries for common context prefixes. However, KV cache fetches can become a bottleneck when netw…