From the 1 of 1 linked paper with an AI index.
1 paper
Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit +2
The paper investigates why reinforcement‑learning‑trained models outperform supervised fine‑tuned models on math reasoning by analyzing their internal representations with linear p…