1 paper
Can Jin, Hongwu Peng, Mingcan Xiang +7
Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-k routing imposes a rigid sparsity pattern that ignores the int…