2 papers
cs.DC2026
ProServe: Unified Multi-Priority Request Scheduling for LLM Serving
Weizhe Huang, Tao Peng, Tongxuan Liu +4
The widespread deployment of large language models (LLMs) for interactive applications necessitates serving systems that can handle thousands of concurrent requests with diverse Se…
cs.DC2026
xLLM Technical Report
Tongxuan Liu, Tao Peng, Peijun Yang +50
We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimi…