1 citations · 1 across the 8 of their papers we have counts for
1 paper · 1 filter
Yun Wang, Lingyun Yang, Senhao Yu +5
Mixture-of-Experts (MoE) architectures scale language models by activating only a subset of specialized expert networks for each input token, thereby reducing the number of floatin…