2 papers
cs.DC2026
Generation Quality-Latency Tradeoff-Aware Inference Offloading for Multimodal LLMs in Cloud-Edge Continuum
Zhongxiao Wang, Yueshen Xu, Xinkui Zhao +2
Beyond pure cloud, some efforts are being made to deploy Large Language Models (LLMs) in edge to accelerate inference response. So the deployment of LLMs in cloud-edge continuum be…
cs.DC2026
Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM Serving
Haoxin Liu, Jiayi Wang, Yueshen Xu +1
As large language models (LLMs) are increasingly deployed with highly heterogeneous workloads, chunked-prefill execution has emerged as a mainstream serving architecture. Balancing…