1 paper · 1 filter
Hao Wang, Rui Li, Lei Sha +1
Existing code reasoning methods primarily supervise final code outputs, ignoring intermediate states, often leading to reward hacking where correct answers are obtained through inc…