AI Safety for Everyone
arXiv:2502.09288 · doi:10.1038/s42256-025-01020-y
Abstract
Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of potential existential threats. However, this framing has three potential drawbacks: it may exclude researchers and practitioners who are committed to AI safety but approach the field from different angles; it could lead the public to mistakenly view AI safety as focused solely on existential scenarios rather than addressing a wide spectrum of safety challenges; and it risks creating resistance to safety measures among those who disagree with predictions of existential AI risks. Through a systematic literature review of primarily peer-reviewed research, we find a vast array of concrete safety work that addresses immediate and practical concerns with current AI systems. This includes crucial areas like adversarial robustness and interpretability, highlighting how AI safety research naturally extends existing technological and systems safety concerns and practices. Our findings suggest the need for an epistemically inclusive and pluralistic conception of AI safety that can accommodate the full range of safety considerations, motivations, and perspectives that currently shape the field.
References in corpus (28)
- Training language models to follow instructions with human feedback
- Visualizing and Understanding Recurrent Networks
- Deep reinforcement learning from human preferences
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- Adversarial Examples Are Not Bugs, They Are Features
- Understanding Adversarial Training: Increasing Local Stability of Neural Nets through Robust Optimization
- On the importance of single directions for generalization
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Zero-Shot Machine Unlearning
- Analyzing Information Leakage of Updates to Natural Language Models
- Unsolved Problems in ML Safety
- Testing Robustness Against Unforeseen Adversaries
- Artificial Intelligence Safety and Cybersecurity: a Timeline of AI Failures
- Toy Models of Superposition
- Sociotechnical Safety Evaluation of Generative AI Systems
- Evaluating Models' Local Decision Boundaries via Contrast Sets
- Bridging the Transparency Gap: What Can Explainable AI Learn From the AI Act?
- When Neurons Fail
- Safe Exploration for Interactive Machine Learning
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
- A robust approach to model-based classification based on trimming and constraints
- A Robust Learning Methodology for Uncertainty-aware Scientific Machine Learning models
- Morality, Machines and the Interpretation Problem: A Value-based, Wittgensteinian Approach to Building Moral Agents
- Concrete Problems in AI Safety, Revisited
- Formal Modelling of Safety Architecture for Responsibility-Aware Autonomous Vehicle via Event-B Refinement
- A Contact-Safe Reinforcement Learning Framework for Contact-Rich Robot Manipulation
- What Would Jiminy Cricket Do? Towards Agents That Behave Morally
- Explicit Explore, Exploit, or Escape (): near-optimal safety-constrained reinforcement learning in polynomial time