1 paper
Junsun Choi, Sam Son, Sunjin Choi +5
Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of LLM serving runtime. This has…