1 paper
Uladzislau Yorsh, Alexander Kovalenko, Vojtěch Vančura +3
In this paper, we propose that the dot product pairwise matching attention layer, which is widely used in Transformer-based models, is redundant for the model performance. Attentio…