1 paper
Zheming Yang, Yuanhao Yang, Chang Zhao +3
With the rapid growth in the number of large language model (LLM) users, it is difficult for bandwidth-constrained cloud servers to simultaneously process massive LLM services in r…