1 paper · 1 filter
Yixiao Chen, Yanyue Xie, Ruining Yang +6
The Mixture of Experts (MoE) architecture is an important method for scaling Large Language Models (LLMs). It increases model capacity while keeping computation cost low. However,…