1 paper · 1 filter
Robin Pan, Raymond Liu, Daniel Fang +2
Mixture-of-Experts (MoE) models enable model scaling while maintaining low inference-time compute by activating only a subset of experts per token. However, conventional routing re…