1 paper
Dong Pan, Bingtao Li, Yongsheng Zheng +2
The sparse Mixture of Experts(MoE) architecture has evolved as a powerful approach for scaling deep learning models to more parameters with comparable computation cost. As an impor…