1 paper
Sagi Ahrac, Noya Hochwald, Mor Geva
Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse onto few experts and auxiliary lo…