7 papers
HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Xin Yuan, Ning Li, Wenchao Xu +3
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challeng…
TrimMoE A communication aware and adaptive depth framework for distributed edge inference
Ning Li, Shuting Bai, Xin Yuan +4
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus…
OrderMoE: An expert similarity driven distributed edge MoE inference
Xin Yuan, Ning Li, Quan Chen +4
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inferenc…
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
Muqing Li, Ning Li, Xin Yuan +4
The proliferation of large language models (LLMs) has driven the adoption of Mixture-of-Experts (MoE) architectures as a promising solution to scale model capacity while controllin…
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
Tuo Zhang, Ning Li, Xin Yuan +4
With the breakthrough progress of large language models (LLMs) in natural language processing and multimodal tasks, efficiently deploying them on resource-constrained edge devices…
FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC
Xin Yuan, Ning Li, Jose Fernan Martinez
Mobile edge computing (MEC) can reduce the latency of cloud computing successfully. However, the edge server may fail due to the hardware of software issues. When the edge server f…