Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
Ilias Kazantzidis, Timothy J. Norman, Yali Du +1
We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward functi…
cs.AI2025
A Comparative User Evaluation of XRL Explanations using Goal Identification
Mark Towers, Yali Du, Christopher Freeman +1
Debugging is a core application of explainable reinforcement learning (XRL) algorithms; however, limited comparative evaluations have been conducted to understand their relative pe…
cs.AI2024
Explaining an Agent's Future Beliefs through Temporally Decomposing Future Reward Estimators
Mark Towers, Yali Du, Christopher Freeman +1
Future reward estimation is a core component of reinforcement learning agents; i.e., Q-value and state-value functions, predicting an agent's sum of future rewards. Their scalar ou…