Managing extreme AI risks amid rapid progress
arXiv:2310.17688 · doi:10.1126/science.adn0117
Abstract
Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, and an irreversible loss of human control over autonomous AI systems. Although researchers have warned of extreme risks from AI, there is a lack of consensus about how exactly such risks arise, and how to manage them. Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems. In this short consensus paper, we describe extreme risks from upcoming, advanced AI systems. Drawing on lessons learned from other safety-critical technologies, we then outline a comprehensive plan combining technical research and development with proactive, adaptive governance mechanisms for a more commensurate preparation.
Published in Science: https://www.science.org/doi/10.1126/science.adn0117
Cited by in corpus (13)
- Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
- Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
- How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
- Crossing the principle-practice gap in AI ethics with ethical problem-solving
- Frontier AI developers need an internal audit function
- Enhancing Trust Through Standards: A Comparative Risk-Impact Framework for Aligning ISO AI Standards with Global Ethical and Regulatory Contexts
- Towards Symbolic XAI -- Explanation Through Human Understandable Logical Relationships Between Features
- Materiality and Risk in the Age of Pervasive AI Sensors
- Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
- A Theory of Intelligences
- The science and practice of proportionality in AI risk evaluations
- Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation
- Resilience to the Flowing Unknown: an Open Set Recognition Framework for Data Streams