17 citations · 38 across the 23 of their papers we have counts for
1 paper · 2 filters
Bo Liu, Leon Guertler, Simon Yu +9
Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approache…