2 papers
cs.DC2026
The KV Cache Working Set: Online Capacity Planning for LLM Inference Systems
Luchang Li, Shuaishuai Wang, Zhao Ruan +2
Prefix caching is critical for efficient large language model (LLM) serving, particularly for agentic workloads that repeatedly invoke the model with a growing conversation and too…
cs.DC2026
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
Luchang Li, Dongfang Li, Bozhao Gong +1
Prefill-Decode (P/D) disaggregation has emerged as a widely adopted optimization strategy for Large Language Model (LLM) inference. However, there currently exists no well-establis…