Showing 2026Show all
3 papers · 1 filter
cs.DC2026
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
Youhe Jiang, Fangcheng Fu, Taiyi Wang +2
Serving Large Language Models (LLMs) can benefit immensely from parallelizing both the model and input requests across multiple devices, but incoming workloads exhibit substantial…
cs.DC2026
Efficient Multi-round LLM Inference over Disaggregated Serving
Wenhao He, Youhe Jiang, Penghao Zhao +4
With the rapid evolution of Large Language Models (LLMs), multi-round workflows, such as autonomous agents and iterative retrieval, have become increasingly prevalent. However, thi…
cs.DC2026
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
Youhe Jiang, Fangcheng Fu, Eiko Yoneki
The rapid growth of large language model (LLM) deployments has made cost-efficient serving systems essential. Recent efforts to enhance system cost-efficiency adopt two main perspe…