1 paper
Qi Wu, Chao Fang, Jiayuan Chen +5
Mixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems wi…