1 citations · 1 across the 12 of their papers we have counts for
4 papers · 2 filters
TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference
Ning Li, Xinyu Wang, Xin Yuan +3
Mixture-of-experts (MoE) models improve capacity with moderate overhead by sparsely activating experts per token. However, deploying MoE across resource-constrained edge servers in…
HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Xin Yuan, Ning Li, Wenchao Xu +2
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challeng…
TrimMoE A communication aware and adaptive depth framework for distributed edge inference
Ning Li, Shuting Bai, Xin Yuan +3
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus…
OrderMoE: An expert similarity driven distributed edge MoE inference
Xin Yuan, Ning Li, Quan Chen +2
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inferenc…