1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Bo Liu, Leon Guertler, Simon Yu +9
Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approache…