1 paper · 1 filter
Fahao Chen, Peng Li, Zicong Hong +2
Mixture-of-Experts (MoE) is an emerging technique for scaling large models with sparse activation. MoE models are typically trained in a distributed manner with an expert paralleli…