1 paper · 1 filter
Yiyang Shen, Lifu Tu, Weiran Wang
Reinforcement Learning (RL) substantially improves the reasoning capabilities of language models, but most existing RL fine-tuning approaches rely entirely on ground-truth verifiab…