1 paper
Xiaoniu Song, Zihang Zhong, Rong Chen +1
The promising applications of large language models are often limited by the constrained GPU memory capacity available on edge devices. Mixture-of-Experts (MoE) models help address…