2 papers
cs.LG2026
Bootstrapping Task Spaces for Self-Improvement
Minqi Jiang, Andrei Lupu, Yoram Bachrach
Progress in many task domains emerges from repeated revisions to previous solution attempts. Training agents that can reliably self-improve over such sequences at inference-time is…
cs.LG2026
Déjà Q: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
Willem Röpke, Samuel Coward, Andrei Lupu +3
Recent advances in reasoning models have yielded impressive results in mathematics and coding. However, most approaches rely on static datasets, which have been suggested to encour…