Normative Conflicts and Shallow AI Alignment
arXiv:2506.04679 · doi:10.1007/s11098-025-02347-3
Abstract
The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing efforts to instill norms such as helpfulness, honesty, and harmlessness in LLMs through fine-tuning based on human preferences, they remain vulnerable to adversarial attacks that exploit conflicts between these norms. I argue that this vulnerability reflects a fundamental limitation of existing alignment methods: they reinforce shallow behavioral dispositions rather than endowing LLMs with a genuine capacity for normative deliberation. Drawing from on research in moral psychology, I show how humans' ability to engage in deliberative reasoning enhances their resilience against similar adversarial tactics. LLMs, by contrast, lack a robust capacity to detect and rationally resolve normative conflicts, leaving them susceptible to manipulation; even recent advances in reasoning-focused LLMs have not addressed this vulnerability. This ``shallow alignment'' problem carries significant implications for AI safety and regulation, suggesting that current approaches are insufficient for mitigating potential harms posed by increasingly capable AI systems.
Published in Philosophical Studies
References in corpus (21)
- Training language models to follow instructions with human feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Scaling Laws for Neural Language Models
- Evaluating Large Language Models Trained on Code
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Adversarial Examples for Evaluating Reading Comprehension Systems
- Improving alignment of dialogue agents via targeted human judgements
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Measuring Faithfulness in Chain-of-Thought Reasoning
- A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates
- Generating Phishing Attacks using ChatGPT
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Large Language Models for Scientific Synthesis, Inference and Explanation
- Will releasing the weights of future large language models grant widespread access to pandemic agents?
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
- Improving Alignment and Robustness with Circuit Breakers
- EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI