2 papers
cs.LG2025
IC-Cache: Efficient Large Language Model Serving via In-context Caching
Yifan Yu, Yu Gan, Nikhil Sarda +7
Large language models (LLMs) have excelled in various applications, yet serving them at scale is challenging due to their substantial resource demands and high latency. Our real-wo…
cs.DC2025
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
Siyuan Chen, Zhipeng Jia, Samira Khan +2
This paper introduces SLOs-Serve, a system designed for serving multi-stage large language model (LLM) requests with application- and stage-specific service level objectives (SLOs)…