1 citations · 1 across the 9 of their papers we have counts for
1 paper · 1 filter
Xucong Wang, Zhe Zhao, Liheng Yu +3
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…