2 papers
cs.LG2026
Deep SPI: Safe Policy Improvement via World Models
Florent Delgrange, Raphael Avalos, Willem Röpke
Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in…
cs.LG2026
Déjà Q: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
Willem Röpke, Samuel Coward, Andrei Lupu +3
Recent advances in reasoning models have yielded impressive results in mathematics and coding. However, most approaches rely on static datasets, which have been suggested to encour…