1 paper
Ziyu Huang, Yangjie Zhou, Chenhao Zhu +15
Mixture-of-Experts (MoE) inference under expert parallelism (EP) turns each MoE layer into a distributed computation with costly dispatch and combine communication. State-of-the-ar…