1 paper
Xiaodong Chen, Mingming Ha, Zhenzhong Lan +2
The Mixture-of-Experts (MoE) architecture has become a predominant paradigm for scaling large language models (LLMs). Despite offering strong performance and computational efficien…