3 papers
cs.DC2026
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
Minxian Xu, Jingfeng Wu, Shengye Song +16
The rapid rise of Large Language Models (LLMs) has revolutionized various artificial intelligence (AI) applications, from natural language processing to code generation. However, t…
cs.DC2026
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
Zizhao Mo, Junlin Chen, Huanle Xu +1
Nowadays, service providers often deploy multiple types of LLM services within shared clusters. While the service colocation improves resource utilization, it introduces significan…
cs.DC2025
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
Zizhao Mo, Jianxiong Liao, Huanle Xu +2
The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end…