1 paper · 1 filter
Junsun Choi, Sam Son, Sunjin Choi +5
Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of LLM serving runtime. This has…