Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
Cameron Pattison, Lorenzo Manuali, Seth Lazar
Safety-trained language models routinely refuse requests for help circumventing rules. But not all rules deserve compliance. When users ask for help evading rules imposed by an ill…
cs.AI2025
Resource Rational Contractualism Should Guide AI Alignment
Sydney Levine, Matija Franklin, Tan Zhi-Xuan +8
AI systems will soon have to navigate human environments and make decisions that affect people and other AI agents whose goals and values diverge. Contractualist alignment proposes…
cs.AI2025
Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs
Daniel Kilov, Caroline Hendy, Secil Yanik Guyot +2
Moral competence is the ability to act in accordance with moral principles. As large language models (LLMs) are increasingly deployed in situations demanding moral competence, ther…