1 paper
Xingkui Zhu, Yiran Guan, Dingkang Liang +3
The sparsely activated mixture of experts (MoE) model presents a promising alternative to traditional densely activated (dense) models, enhancing both quality and computational eff…