1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Xueguang Ma, Qian Liu, Dongfu Jiang +3
Reinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs). Particularly, the "Zero" reinforceme…