Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
Haiying Shen, Tanmoy Sen, Masahiro Tanaka
In Large Language Model (LLM) serving, the KV-cache (KVC) bottleneck causes high tail Time-to-First-Token (TTFT) and Time-Between-Tokens (TBT), impairing user experience, particula…
cs.CL2025
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
Haiying Shen, Tanmoy Sen
In this paper, we consider a mixed-prompt scenario for a large language model (LLM) inference serving system that supports diverse applications with both short prompts and long pro…