1 paper
Svetlana Pavlitska, Haixi Fan, Konstantin Ditschuneit +1
Sparse mixture-of-experts (MoE) layers have been shown to substantially increase model capacity without a proportional increase in computational cost and are widely used in transfo…