Understanding Agent Incentives using Causal Influence Diagrams. Part I: Single Action Settings
arXiv:1902.09980
Abstract
Agents are systems that optimize an objective function in an environment. Together, the goal and the environment induce secondary objectives, incentives. Modeling the agent-environment interaction using causal influence diagrams, we can answer two fundamental questions about an agent's incentives directly from the graph: (1) which nodes can the agent have an incentivize to observe, and (2) which nodes can the agent have an incentivize to control? The answers tell us which information and influence points need extra protection. For example, we may want a classifier for job applications to not use the ethnicity of the candidate, and a reinforcement learning agent not to take direct control of its reward mechanism. Different algorithms and training paradigms can lead to different causal influence diagrams, so our method can be used to identify algorithms with problematic incentives and help in designing algorithms with better incentives.
Mostly superseded by arXiv:2102.01685
References in corpus (7)
- Counterfactual Fairness
- The Measure and Mismeasure of Fairness
- Strong Completeness and Faithfulness in Bayesian Networks
- On Formalizing Fairness in Prediction with Machine Learning
- Bayes-Ball: The Rational Pastime (for Determining Irrelevance and Requisite Information in Belief Networks and Influence Diagrams)
- AI Safety Gridworlds
- Good and safe uses of AI Oracles
Cited by in corpus (11)
- Alignment of Language Agents
- Strategic Classification is Causal Modeling in Disguise
- Incentives for Responsiveness, Instrumental Control and Impact
- Causal Modeling for Fairness in Dynamical Systems
- Algorithms for Causal Reasoning in Probability Trees
- Hidden Incentives for Auto-Induced Distributional Shift
- Causal Analysis of Agent Behavior for AI Safety
- Resolving Spurious Correlations in Causal Models of Environments via Interventions
- Challenges for Using Impact Regularizers to Avoid Negative Side Effects
- Categorizing Wireheading in Partially Embedded Agents
- Intelligence and Unambitiousness Using Algorithmic Information Theory