1 paper
Ke Li, Zheng Yang, Zhongbin Zhou +3
Mixture-of-Experts (MoE) architectures in large language models (LLMs) deliver exceptional performance and reduced inference costs compared to dense LLMs. However, their large para…