5 papers
HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Xin Yuan, Ning Li, Wenchao Xu +3
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challeng…
TrimMoE A communication aware and adaptive depth framework for distributed edge inference
Ning Li, Shuting Bai, Xin Yuan +4
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus…
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
Muqing Li, Ning Li, Xin Yuan +4
The proliferation of large language models (LLMs) has driven the adoption of Mixture-of-Experts (MoE) architectures as a promising solution to scale model capacity while controllin…
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
Tuo Zhang, Ning Li, Xin Yuan +4
With the breakthrough progress of large language models (LLMs) in natural language processing and multimodal tasks, efficiently deploying them on resource-constrained edge devices…
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
Ning Li, Song Guo, Tuo Zhang +5
The powerfulness of LLMs indicates that deploying various LLMs with different scales and architectures on end, edge, and cloud to satisfy different requirements and adaptive hetero…