2 papers
cs.LG2026
LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
Wenkai Chen, Tianshu Li, Wenyong Huang +3
Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth. However, mainstream looped architectures re…
cs.CL2025
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
Yehui Tang, Yichun Yin, Yaoyuan Wang +71
Sparse large language models (LLMs) with Mixture of Experts (MoE) and close to a trillion parameters are dominating the realm of most capable language models. However, the massive…