3 papers
cs.LG2026
It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning
Liam P. H. Mertens, Lucas N. Alegre, Florent Delgrange +3
Time is of the essence when dealing with multiple reward signals and non-linear utility. In this paper we argue that the current main approaches in multi-objective RL (SER and ESR)…
cs.LG2026
Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments
Florent Delgrange
The next generation of autonomous agents must not only learn efficiently but also act reliably and adapt their behavior in open worlds. Standard approaches typically assume fixed t…
cs.LG2025
Deep SPI: Safe Policy Improvement via World Models
Florent Delgrange, Raphael Avalos, Willem Röpke
Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in…