297 citations · 563 across the 10 of their papers we have counts for
1 paper · 2 filters
Weihao Yu, Zihang Jiang, Fei Chen +2
Modern pre-trained language models are mostly built upon backbones stacking self-attention and feed-forward layers in an interleaved order. In this paper, beyond this stereotyped l…