2 citations · 2 across the 8 of their papers we have counts for
1 paper · 1 filter
Gengsheng Li, Jinghan He, Shijie Wang +7
Self-play bootstraps LLM reasoning through an iterative Challenger-Solver loop: the Challenger is trained to generate questions that target the Solver's capabilities, and the Solve…