145 citations · 146 across the 6 of their papers we have counts for
1 paper · 2 filters
Yibo Jin, Tao Wang, Huimin Lin +27
Serving disaggregated large language models (LLMs) over tens of thousands of xPU devices (GPUs or NPUs) with reliable performance faces multiple challenges. 1) Ignoring the diversi…