1 paper
Xucong Wang, Zhe Zhao, Liheng Yu +3
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…