4 papers
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
Joar Skalse, Alessandro Abate
The aim of inverse reinforcement learning (IRL) is to infer an agent's preferences from observing their behaviour. Usually, preferences are modelled as a reward function, , and…
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
Joar Skalse, Alessandro Abate
The aim of Inverse Reinforcement Learning (IRL) is to infer a reward function from a policy . This problem is difficult, for several reasons. First of all, there are typical…
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
Lukas Fluri, Leon Lang, Alessandro Abate +3
In reinforcement learning, specifying reward functions that capture the intended task can be very challenging. Reward learning aims to address this issue by learning the reward fun…
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
David "davidad" Dalrymple, Joar Skalse, Yoshua Bengio +14
Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general in…