3 papers
cs.LG2026
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier
Lorenz Wolf, Connor Watts, Roger Creus Castanyer +4
The limiting resource for training agents via reinforcement learning (RL) is increasingly frontier task supply: valid, solvable tasks just difficult enough to train the current mod…
cs.CR2026
unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning
Geoffrey Bradway, Roger Creus Castanyer, Lorenz Wolf +3
Unix competence is the ability to use shell and operating-system primitives as first-class tools, not merely to write programs through a terminal. Current terminal benchmarks tend…
cs.AI2026
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf +3
We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs. Teachers and students are…