1 paper
Jinze Zhao, Peihao Wang, Junjie Yang +6
Sparse Mixture-of-Experts (SMoE) architectures have gained prominence for their ability to scale neural networks, particularly transformers, without a proportional increase in comp…