Beyond Preferences in AI Alignment
arXiv:2408.16984 · doi:10.1007/s11098-024-02249-w
Abstract
The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and (3) that AI systems should be aligned with the preferences of one or more humans to ensure that they behave safely and in accordance with our values. Whether implicitly followed or explicitly endorsed, these commitments constitute what we term a preferentist approach to AI alignment. In this paper, we characterize and challenge the preferentist approach, describing conceptual and technical alternatives that are ripe for further research. We first survey the limits of rational choice theory as a descriptive model, explaining how preferences fail to capture the thick semantic content of human values, and how utility representations neglect the possible incommensurability of those values. We then critique the normativity of expected utility theory (EUT) for humans and AI, drawing upon arguments showing how rational agents need not comply with EUT, while highlighting how EUT is silent on which preferences are normatively acceptable. Finally, we argue that these limitations motivate a reframing of the targets of AI alignment: Instead of alignment with the preferences of a human user, developer, or humanity-writ-large, AI systems should be aligned with normative standards appropriate to their social roles, such as the role of a general-purpose assistant. Furthermore, these standards should be negotiated and agreed upon by all relevant stakeholders. On this alternative conception of alignment, a multiplicity of AI systems will be able to serve diverse ends, aligned with normative standards that promote mutual benefit and limit harm despite our plural and divergent values.
26 pages (excl. references), 5 figures
References in corpus (43)
- Training language models to follow instructions with human feedback
- CP-nets: A Tool for Representing and Reasoning withConditional Ceteris Paribus Preference Statements
- Artificial Intelligence, Values and Alignment
- Power to the People? Opportunities and Challenges for Participatory AI
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
- On the Acceptability of Arguments in Preference-Based Argumentation
- Jury Learning: Integrating Dissenting Voices into Machine Learning Models
- Fine-tuning language models to find agreement among humans with diverse preferences
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- Collective Constitutional AI: Aligning a Language Model with Public Input
- Faith and Fate: Limits of Transformers on Compositionality
- Reward-rational (implicit) choice: A unifying formalism for reward learning
- On the Planning Abilities of Large Language Models : A Critical Investigation
- Participation in the age of foundation models
- Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
- Online Bayesian Goal Inference for Boundedly-Rational Planning Agents
- An Interval-Valued Utility Theory for Decision Making with Dempster-Shafer Belief Functions
- When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment
- Goal Misgeneralization in Deep Reinforcement Learning
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEval
- Aligning Robot and Human Representations
- Principled Reinforcement Learning with Human Feedback from Pairwise or -wise Comparisons
- On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
- Recognising the importance of preference change: A call for a coordinated multidisciplinary research effort in the age of AI
- Models of human preference for learning reward functions
- Making Intelligence: Ethical Values in IQ and ML Benchmarks
- Language Agents as Digital Representatives in Collective Decision-Making
- The Boltzmann Policy Distribution: Accounting for Systematic Suboptimality in Human Models
- Goals as Reward-Producing Programs
- AI Alignment with Changing and Influenceable Reward Functions
- Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
- Inverse Decision Modeling: Learning Interpretable Representations of Behavior
- Learning and Sustaining Shared Normative Systems via Bayesian Rule Induction in Markov Games
- Modeling the Mistakes of Boundedly Rational Agents Within a Bayesian Theory of Mind
- Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
- Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
- Bayesian Inference of Social Norms as Shared Constraints on Behavior
- A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward Machines
- Modeling Boundedly Rational Agents with Latent Inference Budgets
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
- A density estimation perspective on learning from pairwise human preferences
- Confronting Reward Model Overoptimization with Constrained RLHF
Cited by in corpus (6)
- Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
- Design Considerations for Human Oversight of AI: Insights from Co-Design Workshops and Work Design Theory
- AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
- A Hormetic Approach to the Value-Loading Problem: Preventing the Paperclip Apocalypse?