1 paper
Jinwei Kong, Runqi Meng, Fanyi Wang +4
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited…