1 paper
Cheng Li, Jiexiong Liu, Yixuan Chen +1
Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in compu…