1 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Jing Yu, Shengchao Chen, Yiyun Tan
Reinforcement learning with verifiable rewards has become the dominant recipe for improving large language model reasoning, yet it presumes large human-curated task collections. Ze…