1 paper · 1 filter
Peijun Zhu, Ning Yang, Baoliang Tian +4
Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on…