4 papers
Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning
Tiberiu-Andrei Georgescu, Alexander W. Goodall, Dalal Alrajeh +2
Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent's actions remain compliant with formal specifications. Classical shielding approac…
Safe Reinforcement Learning via Recovery-based Shielding with Gaussian Process Dynamics Models
Alexander W. Goodall, Francesco Belardinelli
Reinforcement learning (RL) is a powerful framework for optimal decision-making and control but often lacks provable guarantees for safety-critical applications. In this paper, we…
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
Alexander W. Goodall, Edwin Hamel-De le Court, Francesco Belardinelli
Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due…
Probabilistic Shielding for Safe Reinforcement Learning
Edwin Hamel-De le Court, Francesco Belardinelli, Alexander W. Goodall
In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximise their reward, must often also behave in a safe manner, including at training time. Thus, much attenti…