1 paper
Yaru Hao, Li Dong, Furu Wei +1
The great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information fro…