164 citations · 183 across the 4 of their papers we have counts for
1 paper · 2 filters
Axel Højmark, Jérémy Scheurer, Evgenia Nitishinskaya +5
Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective. This "reward-seeking" is difficult to measure be…