1 paper · 1 filter
Zhuoran Zhu, Chunyang Zhu, Hao Lin +9
Large-scale Mixture-of-Experts (MoE) models rely on \emph{expert parallelism} for efficient training and inference, which splits experts across devices and necessitates distributed…