1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yibo Jin, Tao Wang, Huimin Lin +27
Serving disaggregated large language models (LLMs) over tens of thousands of xPU devices (GPUs or NPUs) with reliable performance faces multiple challenges. 1) Ignoring the diversi…