1 paper · 1 filter
Dong Pan, Bingtao Li, Yongsheng Zheng +2
The sparse Mixture of Experts(MoE) architecture has evolved as a powerful approach for scaling deep learning models to more parameters with comparable computation cost. As an impor…