1 paper
Tian Wu, Liming Wang, Zijian Wen +5
The emergence of Mixture-of-Experts (MoE) has transformed the scaling of large language models by enabling vast model capacity through sparse activation. Yet, converting these perf…