1 paper
Nakyung Lee, Yeongoon Kim, Minhae Oh +4
Transformer-based self-attention mechanism serves as the core of modern language models, yet it often suffers from localization, where attentions collapse onto a limited subset of…