1 paper · 1 filter
Yibo Jin, Tao Wang, Huimin Lin +27
Serving disaggregated large language models (LLMs) over tens of thousands of xPU devices (GPUs or NPUs) with reliable performance faces multiple challenges. 1) Ignoring the diversi…