Jiayu Zhao, Zihan Teng, Minhao Fan +4
Mixture-of-Experts (MoE) large language models reduce per-token computation through sparse expert activation, but their deployment remains memory-intensive because all expert weigh…