1 citations · 1 across the 9 of their papers we have counts for
1 paper · 1 filter
Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai +6
Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly…