human feedback 1model predictive control 1preference learning 1safe reinforcement learning 1safety-critical AI 1world models 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
Ilias Kazantzidis, Timothy J. Norman, Yali Du +1
The paper introduces DROPJ, a human‑in‑the‑loop approach that uses preferences and justifications collected from users interacting with a learned world model to train a reward mode…
cs.AI2025
A Comparative User Evaluation of XRL Explanations using Goal Identification
Mark Towers, Yali Du, Christopher Freeman +1
Debugging is a core application of explainable reinforcement learning (XRL) algorithms; however, limited comparative evaluations have been conducted to understand their relative pe…