3 papers
cs.LG2025
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
Alexander W. Goodall, Edwin Hamel-De le Court, Francesco Belardinelli
Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due…
stat.ML2025
Probabilistic Shielding for Safe Reinforcement Learning
Edwin Hamel-De le Court, Francesco Belardinelli, Alexander W. Goodall
In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximise their reward, must often also behave in a safe manner, including at training time. Thus, much attenti…
cs.FL2019
Algebraic and Combinatorial Tools for State Complexity : Application to the Star-Xor Problem
Pascal Caron, Edwin Hamel-de le Court, Jean-Gabriel Luque
We investigate the state complexity of the star of symmetrical differences using modifiers and monsters. A monster is an automaton in which every function from states to states is…