849 citations · 890 across the 5 of their papers we have counts for
1 paper · 1 filter
Weihao Yu, Zihang Jiang, Fei Chen +2
Modern pre-trained language models are mostly built upon backbones stacking self-attention and feed-forward layers in an interleaved order. In this paper, beyond this stereotyped l…