Harms from Increasingly Agentic Algorithmic Systems
arXiv:2302.10329 · doi:10.1145/3593013.3594033
Abstract
Research in Fairness, Accountability, Transparency, and Ethics (FATE) has established many sources and forms of algorithmic harm, in domains as diverse as health care, finance, policing, and recommendations. Much work remains to be done to mitigate the serious harms of these systems, particularly those disproportionately affecting marginalized communities. Despite these ongoing harms, new systems are being developed and deployed which threaten the perpetuation of the same harms and the creation of novel ones. In response, the FATE community has emphasized the importance of anticipating harms. Our work focuses on the anticipation of harms from increasingly agentic systems. Rather than providing a definition of agency as a binary property, we identify 4 key characteristics which, particularly in combination, tend to increase the agency of a given algorithmic system: underspecification, directness of impact, goal-directedness, and long-term planning. We also discuss important harms which arise from increasing agency -- notably, these include systemic and/or long-range impacts, often on marginalized stakeholders. We emphasize that recognizing agency of algorithmic systems does not absolve or shift the human responsibility for algorithmic harms. Rather, we use the term agency to highlight the increasingly evident fact that ML systems are not fully under human control. Our work explores increasingly agentic algorithmic systems in three parts. First, we explain the notion of an increase in agency for algorithmic systems in the context of diverse perspectives on agency across disciplines. Second, we argue for the need to anticipate harms from increasingly agentic systems. Third, we discuss important harms from increasingly agentic systems and ways forward for addressing them. We conclude by reflecting on implications of our work for anticipating algorithmic harms from emerging systems.
Accepted at FAccT 2023
References in corpus (26)
- Scaling Laws for Neural Language Models
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- The Fallacy of AI Functionality
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
- Degenerate Feedback Loops in Recommender Systems
- Accountability in an Algorithmic Society: Relationality, Responsibility, and Robustness in Machine Learning
- The Forgotten Margins of AI Ethics
- Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders
- Meaning without reference in large language models
- Teaching language models to support answers with verified quotes
- Scaling Laws for Reward Model Overoptimization
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
- Human-Timescale Adaptation in an Open-Ended Task Space
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
- Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks
- What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring
- Language Models as Agent Models
- AI Ethics for Systemic Issues: A Structural Approach
- Discovering Agents
- Scaling laws for single-agent reinforcement learning
- Path-Specific Objectives for Safer Agent Incentives
- The AI Economist: Optimal Economic Policy Design via Two-level Deep Reinforcement Learning