3 papers
cs.DC2026
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
Junhan Liao, Minxian Xu, Wanyi Zheng +4
To meet strict Service-Level Objectives (SLOs),contemporary Large Language Models (LLMs) decouple the prefill and decoding stages and place them on separate GPUs to mitigate the di…
cs.DC2026
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
Minxian Xu, Junhan Liao, Linfeng Wen +4
Microservice applications are created as loosely coupled application components and they leverage cloud elasticity to reduce costs and increase development speed. However, microser…
cs.DC2025
Cloud Native System for LLM Inference Serving
Minxian Xu, Junhan Liao, Jingfeng Wu +3
Large Language Models (LLMs) are revolutionizing numerous industries, but their substantial computational demands create challenges for efficient deployment, particularly in cloud…