2 papers
cs.LG2026
Deep SPI: Safe Policy Improvement via World Models
Florent Delgrange, Raphael Avalos, Willem Röpke
Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in…
cs.AI2025
Inclusive Fitness as a Key Step Towards More Advanced Social Behaviors in Multi-Agent Reinforcement Learning Settings
Andries Rosseau, Raphaël Avalos, Ann Nowé
The competitive and cooperative forces of natural selection have driven the evolution of intelligence for millions of years, culminating in nature's vast biodiversity and the compl…