1 paper · 1 filter
Hanna Herasimchyk, Robin Labryga, Tomislav Prusina +1
Transformer models systematically favor certain token positions, yet the architectural origins of this position bias remain poorly understood. This bias is closely connected to the…