18 citations · 18 across the 7 of their papers we have counts for
1 paper · 1 filter
Zhiqiang Tan, Maoxin Wang, Sijie Wang +4
It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standa…