1 citations · 2 across the 4 of their papers we have counts for
1 paper · 2 filters
Lei Kang, Jia Li, Mi Tian +1
Sparsely activated Mixture-of-Experts (MoE) models effectively increase the number of parameters while maintaining consistent computational costs per token. However, vanilla MoE mo…