1 paper
Jingkun Liu, Yisong Yue, Max Welling +1
Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When composed across depth, this interacti…