1 paper · 1 filter
Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf +3
We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs. Teachers and students are…