1 paper · 1 filter
Tingfeng Hui, Zhenyu Zhang, Shuohuan Wang +3
Mixture-of-Experts (MoE) shines brightly in large language models (LLMs) and demonstrates outstanding performance in plentiful natural language processing tasks. However, existing…