1 paper
Zhiyuan Zeng, Qipeng Guo, Zhaoye Fei +7
Sparse Mixture of Experts (MoE) models are popular for training large language models due to their computational efficiency. However, the commonly used top-k routing mechanism su…