Artificial Intelligence, Values and Alignment
arXiv:2001.09768 · doi:10.1007/s11023-020-09539-2
Abstract
This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive engagement between people working in both domains. Second, it is important to be clear about the goal of alignment. There are significant differences between AI that aligns with instructions, intentions, revealed preferences, ideal preferences, interests and values. A principle-based approach to AI alignment, which combines these elements in a systematic way, has considerable advantages in this context. Third, the central challenge for theorists is not to identify 'true' moral principles for AI; rather, it is to identify fair principles for alignment, that receive reflective endorsement despite widespread variation in people's moral beliefs. The final part of the paper explores three ways in which fair principles for AI alignment could potentially be identified.
References in corpus (1)
Cited by in corpus (47)
- Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
- Auditing large language models: a three-layered approach
- A Taxonomy of Prompt Modifiers for Text-To-Image Generation
- Ethics-Based Auditing of Automated Decision-Making Systems: Nature, Scope, and Limitations
- Meaningful human control: actionable properties for AI system development
- The Who in XAI: How AI Background Shapes Perceptions of AI Explanations
- Exploring the psychology of LLMs' Moral and Legal Reasoning
- Ethics-Based Auditing of Automated Decision-Making Systems: Intervention Points and Policy Implications
- Identifying and Mitigating the Security Risks of Generative AI
- Mapping Public Perception of Artificial Intelligence: Expectations, Risk-Benefit Tradeoffs, and Value As Determinants for Societal Acceptance
- The Impact of ChatGPT and LLMs on Medical Imaging Stakeholders: Perspectives and Use Cases
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Ethics of generative AI and manipulation: a design-oriented research agenda
- Beyond Preferences in AI Alignment
- Artificial intelligence, rationalization, and the limits of control in the public sector: the case of tax policy optimization
- Artificial virtuous agents in a multiagent tragedy of the commons
- Axes for Sociotechnical Inquiry in AI Research
- Macro Ethics Principles for Responsible AI Systems: Taxonomy and Future Directions
- Artificial Intelligence Can Emulate Human Normative Judgments on Emotional Visual Scenes
- Designing deep neural networks for driver intention recognition
- Cultural Dimensions of AI Perception: Charting Expectations, Risks, Benefits, Tradeoffs, and Value in Germany and China
- Human participants in AI research: Ethics and transparency in practice
- Normative Conflicts and Shallow AI Alignment
- C3AI: Crafting and Evaluating Constitutions for Constitutional AI
- Beyond Algorethics: Addressing the Ethical and Anthropological Challenges of AI Recommender Systems
- A General Approach for Computing a Consensus in Group Decision Making That Integrates Multiple Ethical Principles
- Perception Gaps in Risk, Benefit, and Value Between Experts and Public Challenge Socially Accepted AI
- Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
- A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
- Aligning Human Intent from Imperfect Demonstrations with Confidence-based Inverse soft-Q Learning
- Assessment of cognitive characteristics in intelligent systems and predictive ability
- Artificial Intelligence in Deliberation: The AI Penalty and the Emergence of a New Deliberative Divide
- Are Large Language Models Aligned with People's Social Intuitions for Human-Robot Interactions?
- Developing and Evaluating a Design Method for Positive Artificial Intelligence
- Case Law Grounding: Using Precedents to Align Decision-Making for Humans and AI
- Towards Interactive Reinforcement Learning with Intrinsic Feedback
- Several Issues Regarding Data Governance in AGI
- Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values
- TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
- Medical Knowledge Integration into Reinforcement Learning Algorithms for Dynamic Treatment Regimes
- Responsible AI: The Good, The Bad, The AI
- AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
- Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI
- Learning the Value Systems of Agents with Preference-based and Inverse Reinforcement Learning
- Detecting underdetermination in parameterized quantum circuits
- A Fair and Ethical Healthcare Artificial Intelligence System for Monitoring Driver Behavior and Preventing Road Accidents
- When Robots Say No: The Empathic Ethical Disobedience Benchmark