4 papers
OrderMoE: An expert similarity driven distributed edge MoE inference
Xin Yuan, Ning Li, Quan Chen +4
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inferenc…
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
Muqing Li, Ning Li, Xin Yuan +4
The proliferation of large language models (LLMs) has driven the adoption of Mixture-of-Experts (MoE) architectures as a promising solution to scale model capacity while controllin…
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
Tuo Zhang, Ning Li, Xin Yuan +4
With the breakthrough progress of large language models (LLMs) in natural language processing and multimodal tasks, efficiently deploying them on resource-constrained edge devices…
FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC
Xin Yuan, Ning Li, Jose Fernan Martinez
Mobile edge computing (MEC) can reduce the latency of cloud computing successfully. However, the edge server may fail due to the hardware of software issues. When the edge server f…