1 paper
Longfei Yun, Yonghao Zhuang, Yao Fu +2
Mixture-of-Expert (MoE) based large language models (LLMs), such as the recent Mixtral and DeepSeek-MoE, have shown great promise in scaling model size without suffering from the q…