1 paper
Fan Mo, Yuxuan Han, Geng Zhang +2
Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse ac…