1 paper · 1 filter
Zizhao Mo, Jianxiong Liao, Huanle Xu +2
The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end…