1 paper
Shuhan Huang, Naifan Zhang, Yuanbo Tang +2
Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition expert select…