3 papers
cs.LG2026
Exploring Starts Are Not Enough: Counterexamples and a Fix for Monte Carlo Exploring Starts
Octave Oliviers, Glenn Vinnicombe
The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting. We investigated the converg…
cs.LG2026
Convergence of Monte Carlo Optimistic Policy Iteration: Beyond Uniform State-Action Updates
Octave Oliviers, Glenn Vinnicombe
The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question. When the model of the environment is unknown, as is common in practi…
cs.LG2025
Deep Learning Agents Trained For Avoidance Behave Like Hawks And Doves
Aryaman Reddi
We present heuristically optimal strategies expressed by deep learning agents playing a simple avoidance game. We analyse the learning and behaviour of two agents within a symmetri…