1 paper
Ning Li, Shuting Bai, Xin Yuan +4
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus…