2 citations · 4 across the 12 of their papers we have counts for
1 paper · 2 filters
Sara Rajaee, Kumar Pratik, Gabriele Cesa +1
The most promising recent methods for AI reasoning require applying variants of reinforcement learning (RL) either on rolled out trajectories from the LLMs, even for the step-wise…