1 paper
Luca Zhou, Sajel Shah, Emanuele Rodolà +1
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with…