117 citations · 184 across the 13 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.AI2023
Honesty Is the Best Policy: Defining and Mitigating AI Deception
Francis Rhys Ward, Francesco Belardinelli, Francesca Toni +1
Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (…
cs.AI2023
Characterising Decision Theories with Mechanised Causal Graphs
Matt MacDermott, Tom Everitt, Francesco Belardinelli
How should my own decisions affect my beliefs about the outcomes I expect to achieve? If taking a certain action makes me view myself as a certain type of person, it might affect h…