3 papers
cs.CL2025
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
Haiying Shen, Tanmoy Sen, Masahiro Tanaka
In Large Language Model (LLM) serving, the KV-cache (KVC) bottleneck causes high tail Time-to-First-Token (TTFT) and Time-Between-Tokens (TBT), impairing user experience, particula…
cs.DC2025
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
Haiying Shen, Tanmoy Sen
As Large Language Models (LLMs) continue to grow, reducing costs and alleviating GPU demands has become increasingly critical. However, existing schedulers primarily target either…
cs.CL2025
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
Haiying Shen, Tanmoy Sen
In this paper, we consider a mixed-prompt scenario for a large language model (LLM) inference serving system that supports diverse applications with both short prompts and long pro…