1.8k citations
- Zhejiang UniversityCN54 papers
- Peking UniversityCN29 papers
- Alibaba Group (United States)US27 papers
- Tsinghua UniversityCN25 papers
- Shanghai Jiao Tong UniversityCN20 papers
- University of Science and Technology of ChinaCN18 papers
- Chinese Academy of SciencesCN15 papers
- Wuhan UniversityCN14 papers
- Nanyang Technological UniversitySG12 papers
- University of Chinese Academy of SciencesCN11 papers
- Huazhong University of Science and TechnologyCN10 papers
- Hong Kong University of Science and TechnologyHK9 papers
Showing 2025 · cs.DCShow all
2 papers · 2 filters
cs.DC2025★ 1 cited
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
Yiyuan He, Minxian Xu, Jingfeng Wu +7
Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, wh…
cs.DC2025
STAR: Decode-Phase Rescheduling for LLM Inference
Zhibin Wang, Zetao Hong, Xue Li +8
Large Language Model (LLM) inference has emerged as a fundamental paradigm, however, variations in output length cause severe workload imbalance in the decode phase, particularly f…