19 citations · 35 across the 15 of their papers we have counts for
1 paper · 2 filters
Lishui Fan, Yu Zhang, Mouxiang Chen +1
In practice, rigorous reasoning is often a key driver of correct code, while Reinforcement Learning (RL) for code generation often neglects optimizing reasoning quality. Bringing p…