1 paper
Jing Yu, Shengchao Chen, Yiyun Tan
Reinforcement learning with verifiable rewards has become the dominant recipe for improving large language model reasoning, yet it presumes large human-curated task collections. Ze…