1 paper
Tianshu Zhu, Wenyu Zhang, Xiaoying Zuo +8
Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrasti…