1 paper
Xiongwei Zhu, Xiaojian Liao, Tianyang Jiang +3
Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining high model capacity. However, in…