1 paper · 1 filter
Zeliang Zhang, Xiaodong Liu, Hao Cheng +2
By increasing model parameters but activating them sparsely when performing a task, the use of Mixture-of-Experts (MoE) architecture significantly improves the performance of Large…