2 papers
cs.AI2026
DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
Can Jin, Hongwu Peng, Mingcan Xiang +7
Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top- routing imposes a rigid sparsity pattern that ignores the int…
cs.LG2026
Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models
Hongwu Peng, Ohiremen Dibua, Yuanjun Xiong +3
We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer blocks. Existing tools such as $Î…