4 papers
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
Omar Basit, Yunzhao Liu, Z. Jonny Kong +1
Prefill/decode disaggregation is increasingly adopted in LLM serving to improve the latency-throughput tradeoff and meet strict TTFT and TPOT SLOs. However, LLM inference remains e…
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
Yunzhao Liu, Qiang Xu, Y. Charlie Hu
Efficient LLM inference is critical for real-world applications, especially within heterogeneous GPU clusters commonly found in organizations and on-premise datacenters as GPU arch…
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
Z. Jonny Kong, Qiang Xu, Y. Charlie Hu
With the rapid innovation of GPUs, heterogeneous GPU clusters in both public clouds and on-premise data centers have become increasingly commonplace. In this paper, we demonstrate…
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
Zongpu Zhang, Pranab Dash, Y. Charlie Hu +3
Large Language Models (LLMs) are increasingly being integrated into various applications and services running on billions of mobile devices. However, deploying LLMs on resource-lim…