1 paper
Shuhuai Li, Jianghao Lin, Dongdong Ge +1
Mixture-of-Experts (MoE) models enable scalable performance but face severe memory constraints on edge devices. Existing offloading strategies struggle with I/O bottlenecks due to…