1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.DC2025
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
Junhan Liao, Minxian Xu, Wanyi Zheng +4
To meet strict Service-Level Objectives (SLOs),contemporary Large Language Models (LLMs) decouple the prefill and decoding stages and place them on separate GPUs to mitigate the di…
cs.DC2025
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
Wanyi Zheng, Minxian Xu, Shengye Song +1
Large language models (LLMs) have become increasingly popular in various areas, traditional business gradually shifting from rule-based systems to LLM-based solutions. However, the…
cs.DC2024★ 1 cited
UELLM: A Unified and Efficient Approach for LLM Inference Serving
Yiyuan He, Minxian Xu, Jingfeng Wu +3
In the context of Machine Learning as a Service (MLaaS) clouds, the extensive use of Large Language Models (LLMs) often requires efficient management of significant query loads. Wh…