4 citations · 5 across the 4 of their papers we have counts for
1 paper · 1 filter
Sophia Xiao Pu, Zhaotian Weng, Chengzhi Liu +4
Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent systems report strong reasoning…