1 paper
Huanyu Wang, Ziyu Xia, Zhuoming Chen +1
Large language model (LLM) services are mostly centralized, leading to scalability bottlenecks and underutilization of substantial scattered GPU resources. While decentralization o…