1 paper
Xingyu Xiang, Raj Joshi, Yuhan Liu +6
Distributed prefix caching accelerates long-context LLM serving by reusing KV cache entries for common context prefixes. However, KV cache fetches can become a bottleneck when netw…