1 paper
Yuval Ran-Milo, Hila Ofek, Shahar Mendel
Transformers commonly exhibit an attention sink: disproportionately high attention to the first position. We study this behavior in GPT-2-style models with learned query biases and…