2 papers
cs.LG2026
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
Chengyi Nie, Nian Si, Zijie Zhou
The rapid adoption of large language models (LLMs) has created significant challenges for efficient inference at scale. Unlike traditional workloads, LLM inference is constrained b…
cs.DC2024
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
Chengyi Nie, Rodrigo Fonseca, Zhenhua Liu
The demand for large language model (LLM) inference is gradually dominating the artificial intelligence workloads. Therefore, there is an urgent need for cost-efficient inference s…