collaborators

7 papers

cs.NI2026

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference

Xin Yuan, Ning Li, Wenchao Xu +3

Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challeng…

cs.NI2026

TrimMoE A communication aware and adaptive depth framework for distributed edge inference

Ning Li, Shuting Bai, Xin Yuan +4

Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus…

cs.NI2026

OrderMoE: An expert similarity driven distributed edge MoE inference

Xin Yuan, Ning Li, Quan Chen +4

Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inferenc…

cs.NI2025

CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge

Muqing Li, Ning Li, Xin Yuan +4

The proliferation of large language models (LLMs) has driven the adoption of Mixture-of-Experts (MoE) architectures as a promising solution to scale model capacity while controllin…

cs.LG2025

Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative

Tuo Zhang, Ning Li, Xin Yuan +4

With the breakthrough progress of large language models (LLMs) in natural language processing and multimodal tasks, efficiently deploying them on resource-constrained edge devices…

cs.DC2025

FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC

Xin Yuan, Ning Li, Jose Fernan Martinez

Mobile edge computing (MEC) can reduce the latency of cloud computing successfully. However, the edge server may fail due to the hardware of software issues. When the edge server f…