3 citations · 7 across the 4 of their papers we have counts for
6 papers
Path-Specific Objectives for Safer Agent Incentives
Sebastian Farquhar, Ryan Carey, Tom Everitt
We present a general framework for training safe agents whose naive incentives are unsafe. As an example, manipulative or deceptive behaviour can improve rewards but should be avoi…
Too Big to Fail? Active Few-Shot Learning Guided Logic Synthesis
Animesh Basak Chowdhury, Benjamin Tan, Ryan Carey +3
Generating sub-optimal synthesis transformation sequences ("synthesis recipe") is an important problem in logic synthesis. Manually crafted synthesis recipes have poor quality. Sta…
A Complete Criterion for Value of Information in Soluble Influence Diagrams
Chris van Merwijk, Ryan Carey, Tom Everitt
Influence diagrams have recently been used to analyse the safety and fairness properties of AI systems. A key building block for this analysis is a graphical criterion for value of…
Why Fair Labels Can Yield Unfair Predictions: Graphical Conditions for Introduced Unfairness
Carolyn Ashurst, Ryan Carey, Silvia Chiappa +1
In addition to reproducing discriminatory relationships in the training data, machine learning systems can also introduce or amplify discriminatory effects. We refer to this as int…
Agent Incentives: A Causal Perspective
Tom Everitt, Ryan Carey, Eric Langlois +2
We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a n…
(When) Is Truth-telling Favored in AI Debate?
Vojtěch Kovařík, Ryan Carey
For some problems, humans may not be able to accurately judge the goodness of AI-proposed solutions. Irving et al. (2018) propose that in such cases, we may use a debate between tw…