1 paper · 1 filter
Zheyue Tan, Zhiyuan Li, Tao Yuan +13
Mixture-of-Experts (MoE) architectures have emerged as a promising approach to scale Large Language Models (LLMs). MoE boosts the efficiency by activating a subset of experts per t…