1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2024
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
Joar Skalse, Alessandro Abate
Inverse reinforcement learning (IRL) aims to infer an agent's preferences (represented as a reward function ) from their behaviour (represented as a policy ). To do this, we…
cs.AI2024
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
Joar Skalse, Alessandro Abate
In this paper, we study the expressivity of scalar, Markovian reward functions in Reinforcement Learning (RL), and identify several limitations to what they can express. Specifical…
cs.LG2023★ 1 cited
Goodhart's Law in Reinforcement Learning
Jacek Karwowski, Oliver Hayman, Xingjian Bai +3
Implementing a reward function that perfectly captures a complex task in the real world is impractical. As a result, it is often appropriate to think of the reward function as a pr…