1 paper
Toby Simonds, Kevin Lopez, Akira Yoshiyama +1
Large language models can generate solutions to complex problems, but training them with reinforcement learning typically requires verifiable rewards that are expensive to create a…