12 citations · 52 across the 37 of their papers we have counts for
8 papers · 1 filter
TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference
Ning Li, Xinyu Wang, Xin Yuan +3
Mixture-of-experts (MoE) models improve capacity with moderate overhead by sparsely activating experts per token. However, deploying MoE across resource-constrained edge servers in…
HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Xin Yuan, Ning Li, Wenchao Xu +2
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challeng…
TrimMoE A communication aware and adaptive depth framework for distributed edge inference
Ning Li, Shuting Bai, Xin Yuan +3
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus…
OrderMoE: An expert similarity driven distributed edge MoE inference
Xin Yuan, Ning Li, Quan Chen +2
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inferenc…
Graph Neural Network-Based Multicast Routing for On-Demand Streaming Services in 6G Networks
Xiucheng Wang, Zien Wang, Nan Cheng +3
The increase of bandwidth-intensive applications in sixth-generation (6G) wireless networks, such as real-time volumetric streaming and multi-sensory extended reality, demands inte…
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
Muqing Li, Ning Li, Xin Yuan +4
The proliferation of large language models (LLMs) has driven the adoption of Mixture-of-Experts (MoE) architectures as a promising solution to scale model capacity while controllin…