2 papers
cs.LG2026
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
Yuval Ran-Milo, Hila Ofek, Shahar Mendel
Transformers commonly exhibit an attention sink: disproportionately high attention to the first position. We study this behavior in GPT-2-style models with learned query biases and…
cs.LG2026
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
Yuval Ran-Milo, Yotam Alexander, Shahar Mendel +1
Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermediate reasoning steps (Chain-of-Thought…