1 paper
Zewen Jin, Shengnan Wang, Jiaan Zhu +5
The Mixture-of-Experts (MoE) structure scales the Transformer-based large language models (LLMs) and improves their performance with only the sub-linear increase in computation res…