2 papers
cs.DC2026
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
Omar Basit, Yunzhao Liu, Z. Jonny Kong +1
Prefill/decode disaggregation is increasingly adopted in LLM serving to improve the latency-throughput tradeoff and meet strict TTFT and TPOT SLOs. However, LLM inference remains e…
cs.DC2025
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
Yunzhao Liu, Qiang Xu, Y. Charlie Hu
Efficient LLM inference is critical for real-world applications, especially within heterogeneous GPU clusters commonly found in organizations and on-premise datacenters as GPU arch…