Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
SHARD: Safe and Helpful Alignment via Self-Reframing Distillation
Viswonathan Manoranjan, Amogh Gupta, Anvesh Rao Vijjini +2
Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational ne…
cs.CL2025
Are language models rational? The case of coherence norms and belief revision
Thomas Hofweber, Peter Hase, Elias Stengel-Eskin +1
Do norms of rationality apply to machine learning models, in particular language models? In this paper we investigate this question by focusing on a special subset of rational norm…
cs.CL2024
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
Peter Hase, Thomas Hofweber, Xiang Zhou +2
The model editing problem concerns how language models should learn new facts about the world over time. While empirical research on model editing has drawn widespread attention, t…