1 paper
Haochen Huang, Shuzhang Zhong, Shengxuan Qiu +8
Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. However,…