1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Manxi Sun, Wei Liu, Jian Luan +2
The Sparsely-Activated Mixture-of-Experts (MoE) has gained increasing popularity for scaling up large language models (LLMs) without exploding computational costs. Despite its succ…