6 citations · 6 across the 6 of their papers we have counts for
1 paper · 1 filter
Chang Tian, Matthew B. Blaschko, Mingzhe Xing +3
Reinforcement learning (RL) has become a key technique for enhancing the reasoning abilities of large language models (LLMs), with policy-gradient algorithms dominating the post-tr…