1 paper
Ziyi Han, Xutong Liu, Ruiting Zhou +2
Sparse Mixture of Experts (SMoE) has become a preferred architecture for scaling Transformer capacity without increasing computational cost, as it activates only a small subset of…