10 citations · 17 across the 26 of their papers we have counts for
1 paper · 2 filters
Houyi Li, Ka Man Lo, Shijie Xuyang +7
Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense…