3 papers
cs.DC2025
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
Yunzhao Liu, Qiang Xu, Y. Charlie Hu
Efficient LLM inference is critical for real-world applications, especially within heterogeneous GPU clusters commonly found in organizations and on-premise datacenters as GPU arch…
cs.DC2025
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
Z. Jonny Kong, Qiang Xu, Y. Charlie Hu
With the rapid innovation of GPUs, heterogeneous GPU clusters in both public clouds and on-premise data centers have become increasingly commonplace. In this paper, we demonstrate…
cs.OS2025
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
Zongpu Zhang, Pranab Dash, Y. Charlie Hu +3
Large Language Models (LLMs) are increasingly being integrated into various applications and services running on billions of mobile devices. However, deploying LLMs on resource-lim…