2 papers
cs.LG2026
Exploring Starts Are Not Enough: Counterexamples and a Fix for Monte Carlo Exploring Starts
Octave Oliviers, Glenn Vinnicombe
The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting. We investigated the converg…
cs.LG2026
Convergence of Monte Carlo Optimistic Policy Iteration: Beyond Uniform State-Action Updates
Octave Oliviers, Glenn Vinnicombe
The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question. When the model of the environment is unknown, as is common in practi…