1 paper · 1 filter
Bowen Liu, Zhi Wu, Runquan Xie +2
Reinforcement Learning from Verifiable Rewards (RLVR) is bottlenecked by data: existing synthesis pipelines rely on expert-written code or fixed templates, confining growth to inst…