1 paper
Hongbo Li, Qinhang Wu, Sen Lin +2
Mixture-of-Experts (MoE) models improve transformer efficiency but lack a unified theoretical explanation, especially when both feed-forward and attention layers are allowed to spe…