Unsolved Problems in ML Safety
arXiv:2109.13916
Abstract
Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the technical problems that the field needs to address. We present four problems ready for research, namely withstanding hazards ("Robustness"), identifying hazards ("Monitoring"), reducing inherent model hazards ("Alignment"), and reducing systemic hazards ("Systemic Safety"). Throughout, we clarify each problem's motivation and provide concrete research directions.
Position Paper
References in corpus (30)
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- On the Opportunities and Risks of Foundation Models
- Equality of Opportunity in Supervised Learning
- On Calibration of Modern Neural Networks
- Evaluating Large Language Models Trained on Code
- CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning
- Theoretically Principled Trade-off between Robustness and Accuracy
- Deep Learning Scaling is Predictable, Empirically
- Deep Anomaly Detection with Outlier Exposure
- Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration
- Scaling Laws for Autoregressive Generative Modeling
- Fixing Data Augmentation to Improve Adversarial Robustness
- SOREL-20M: A Large Scale Benchmark Dataset for Malicious PE Detection
- Unsupervised Translation of Programming Languages
- Detecting and Characterizing Lateral Phishing at Scale
- Test-Time Adaptation to Distribution Shift by Confidence Maximization and Input Transformation
- Alignment of Language Agents
- Program Synthesis with Large Language Models
- AI Research Considerations for Human Existential Safety (ARCHES)
- What are you optimizing for? Aligning Recommender Systems with Human Values
- Handcrafted Backdoors in Deep Neural Networks
- Poisoning and Backdooring Contrastive Learning
- Open Problems in Cooperative AI
- Conservative Objective Models for Effective Offline Model-Based Optimization
- Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
- Hidden Incentives for Auto-Induced Distributional Shift
- Out-of-Distribution Dynamics Detection: RL-Relevant Benchmarks and Results
- What Would Jiminy Cricket Do? Towards Agents That Behave Morally
- Augmenting Decision Making via Interactive What-If Analysis
Cited by in corpus (11)
- On the Opportunities and Risks of Foundation Models
- Unleashing the potential of prompt engineering for large language models
- Worldwide AI Ethics: a review of 200 guidelines and recommendations for AI governance
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- Deception Abilities Emerged in Large Language Models
- AI Safety for Everyone
- A General Language Assistant as a Laboratory for Alignment
- A Unified Survey on Anomaly, Novelty, Open-Set, and Out-of-Distribution Detection: Solutions and Future Challenges
- Overparameterized Linear Regression under Adversarial Attacks
- Cycle Consistency-based Uncertainty Quantification of Neural Networks in Inverse Imaging Problems
- Certified Adversarial Defenses Meet Out-of-Distribution Corruptions: Benchmarking Robustness and Simple Baselines