4 papers · 1 filter
Generation Quality-Latency Tradeoff-Aware Inference Offloading for Multimodal LLMs in Cloud-Edge Continuum
Zhongxiao Wang, Yueshen Xu, Xinkui Zhao +2
Beyond pure cloud, some efforts are being made to deploy Large Language Models (LLMs) in edge to accelerate inference response. So the deployment of LLMs in cloud-edge continuum be…
Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM Serving
Haoxin Liu, Jiayi Wang, Yueshen Xu +1
As large language models (LLMs) are increasingly deployed with highly heterogeneous workloads, chunked-prefill execution has emerged as a mainstream serving architecture. Balancing…
MUSE: A Heterogeneity-Aware Multimedia Search Engine for Mobile SoCs
Xinkui Zhao, Qingyu Ma, Yifan Zhang +8
On-device multimedia retrieval is vital for smartphones, enabling applications like cross-modal semantic search and multimodal personal AI agents. However, realizing efficient retr…
Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
Chen Chen, Xinkui Zhao, Guanjie Cheng +3
Interconnection is crucial for computing systems. However, the current interconnection performance between processors and devices, such as memory devices and accelerators, signific…