1 paper · 1 filter
Cheng Li, Jiexiong Liu, Yixuan Chen +1
Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in compu…