1 paper · 1 filter
Zhiheng Xi, Wenxiang Chen, Boyang Hong +18
In this paper, we propose R3: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the bene…