1 paper
Shabari S Nair, Krishanu Saini
Prefix caching can reduce LLM inference latency by reusing KV caches across requests with shared prompts, but cluster-scale reuse is challenging because caches are partitioned acro…