3 papers
cs.LG2025
PAC Apprenticeship Learning with Bayesian Active Inverse Reinforcement Learning
Ondrej Bajgar, Dewi S. W. Gould, Jonathon Liu +3
As AI systems become increasingly autonomous, reliably aligning their decision-making with human preferences is essential. Inverse reinforcement learning (IRL) offers a promising a…
cs.CY2025
Position: Ensuring mutual privacy is necessary for effective external evaluation of proprietary AI systems
Ben Bucknall, Robert F. Trager, Michael A. Osborne
The external evaluation of AI systems is increasingly recognised as a crucial approach for understanding their potential risks. However, facilitating external evaluation in practic…
cs.LG2024
Walking the Values in Bayesian Inverse Reinforcement Learning
Ondrej Bajgar, Alessandro Abate, Konstantinos Gatsis +1
The goal of Bayesian inverse reinforcement learning (IRL) is recovering a posterior distribution over reward functions using a set of demonstrations from an expert optimizing for a…