1 paper
Junzhuo Li, Peijie Jiang, Changxin Tian +3
This paper presents a novel extension of neural scaling laws to Mixture-of-Experts (MoE) models, focusing on the optimal allocation of compute between expert and attention sub-laye…