1 paper
Aozhe Wang, Yuchen Yan, Nan Zhou +5
Reinforcement learning for code generation relies on verifiable rewards from unit test pass rates. Yet high-quality test suites are scarce, existing datasets offer limited coverage…