117 citations · 184 across the 11 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.AI2022★ 3 cited
Path-Specific Objectives for Safer Agent Incentives
Sebastian Farquhar, Ryan Carey, Tom Everitt
We present a general framework for training safe agents whose naive incentives are unsafe. As an example, manipulative or deceptive behaviour can improve rewards but should be avoi…
cs.AI2022
A Complete Criterion for Value of Information in Soluble Influence Diagrams
Chris van Merwijk, Ryan Carey, Tom Everitt
Influence diagrams have recently been used to analyse the safety and fairness properties of AI systems. A key building block for this analysis is a graphical criterion for value of…