6 papers
Online Conformal Prediction Beyond Feedback
Joar Skalse, Edoardo Pona, Osvaldo Simeone +1
Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provides theoretically principled…
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
Lukas Fluri, Leon Lang, Alessandro Abate +3
In reinforcement learning, specifying reward functions that capture the intended task can be very challenging. Reward learning aims to address this issue by learning the reward fun…
Defining and Characterizing Reward Hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov +1
We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward fu…
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
Joar Skalse, Alessandro Abate
The aim of inverse reinforcement learning (IRL) is to infer an agent's preferences from observing their behaviour. Usually, preferences are modelled as a reward function, , and…
STARC: A General Framework For Quantifying Differences Between Reward Functions
Joar Skalse, Lucy Farnik, Sumeet Ramesh Motwani +3
In order to solve a task using reinforcement learning, it is necessary to first formalise the goal of that task as a reward function. However, for many real-world tasks, it is very…
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
Joar Skalse, Alessandro Abate
The aim of Inverse Reinforcement Learning (IRL) is to infer a reward function from a policy . This problem is difficult, for several reasons. First of all, there are typica…