1 paper · 1 filter
Yuan Xie, Shaohan Huang, Tianyu Chen +1
Sparsely Mixture of Experts (MoE) has received great interest due to its promising scaling capability with affordable computational overhead. MoE converts dense layers into sparse…