Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
WWW.Serve: Interconnecting Global LLM Services through Decentralization
Huanyu Wang, Ziyu Xia, Zhuoming Chen +1
Large language model (LLM) services are mostly centralized, leading to scalability bottlenecks and underutilization of substantial scattered GPU resources. While decentralization o…
cs.DC2024
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
Youhe Jiang, Ran Yan, Xiaozhe Yao +3
Serving generative inference of the large language model is a crucial component of contemporary AI applications. This paper focuses on deploying such services in a heterogeneous an…