4 citations · 4 across the 3 of their papers we have counts for
1 paper · 1 filter
Andrew Lee, Lihao Sun, Chris Wendler +2
How do reasoning models verify their own answers? We study this question by training a model using DeepSeek R1's recipe on the CountDown task. We leverage the fact that preference…