1 paper · 1 filter
Nakyung Lee, Yeongoon Kim, Minhae Oh +4
Transformer-based self-attention mechanism serves as the core of modern language models, yet it often suffers from localization, where attentions collapse onto a limited subset of…