9 citations · 9 across the 2 of their papers we have counts for
1 paper · 1 filter
Longfei Yun, Yonghao Zhuang, Yao Fu +2
Mixture-of-Expert (MoE) based large language models (LLMs), such as the recent Mixtral and DeepSeek-MoE, have shown great promise in scaling model size without suffering from the q…