1 paper
Bowen Zhou, Jinrui Jia, Wenhao He +2
The Mixture of Experts (MoE) models are emerging as the latest paradigm for Large Language Models (LLMs). However, due to memory constraints, MoE models with billions or even trill…