2 papers
cs.DC2026
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
Weiye Wang, Chen Chen, Junxue Zhang +7
Distributed prefix caching has become a core technique for efficient LLM serving. However, for long-context requests with high cache hit ratios, retrieving reusable KVCache blocks…
cs.DC2025
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
Yifei Liu, Zuo Gan, Zhenghao Gan +8
Applications based on Large Language Models (LLMs) contains a series of tasks to address real-world problems with boosted capability, which have dynamic demand volumes on diverse b…