1 citations · 1 across the 4 of their papers we have counts for
3 papers · 1 filter
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
Cameron Pattison, Lorenzo Manuali, Seth Lazar
Safety-trained language models routinely refuse requests for help circumventing rules. But not all rules deserve compliance. When users ask for help evading rules imposed by an ill…
Resource Rational Contractualism Should Guide AI Alignment
Sydney Levine, Matija Franklin, Tan Zhi-Xuan +8
AI systems will soon have to navigate human environments and make decisions that affect people and other AI agents whose goals and values diverge. Contractualist alignment proposes…
Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs
Daniel Kilov, Caroline Hendy, Secil Yanik Guyot +2
Moral competence is the ability to act in accordance with moral principles. As large language models (LLMs) are increasingly deployed in situations demanding moral competence, ther…