1 paper · 1 filter
Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang +3
Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewar…