1 citations · 1 across the 5 of their papers we have counts for
1 paper · 2 filters
Jama Hussein Mohamud, Drew Wagner, Mirco Ravanelli
Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to exper…