1 paper
Ying Yuan, Pengfei Zuo, Bo Wang +3
In LLM serving, reusing the KV cache of prompts across requests is critical for reducing TTFT and serving costs. Cache-affinity scheduling, which co-locates requests with the same…