1 paper · 1 filter
Nizar Islah, Istabrak Abbes, Irina Rish +2
When post-trained language models fail on reasoning problems, the common test-time-scaling response is to spend more compute on additional attempts, and the failed traces play no f…