3 papers
cs.SE2026
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
Yuanhao Li, Hongbo Wang, Xuhong Chen +2
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterf…
cs.AI2026
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
Yuanhao Li, Hongbo Wang, Xiaotang Shang +3
Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually fix bugs. We present BoostAPR…
cs.CL2025
Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
Qiyuan Liu, Hao Xu, Xuhong Chen +3
Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune LLMs during reinforcement learni…