1 paper
Chengze Du, Zhiwei Yu, Heng Xu +3
The rapid growth of large language model (LLM) services imposes increasing demands on distributed GPU inference infrastructure. Most existing scheduling systems follow a reactive p…