1 paper · 1 filter
Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2
Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical…