1 paper
Rongzhi Li, Ruogu Du, Zefang Chu +8
Serving Large Language Models (LLMs) is a GPU-intensive task where traditional autoscalers fall short, particularly for modern Prefill-Decode (P/D) disaggregated architectures. Thi…