1 paper
Mingyu Cao, Gen Li, Jie Ji +6
Mixture-of-Experts (MoE) has garnered significant attention for its ability to scale up neural networks while utilizing the same or even fewer active parameters. However, MoE does…