2 papers
cs.SE2026
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
Yuanhao Li, Hongbo Wang, Xuhong Chen +2
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterf…
cs.AI2026
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
Yuanhao Li, Hongbo Wang, Xiaotang Shang +3
Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually fix bugs. We present BoostAPR…