7 papers · 1 filter
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
Qinan Yu, Alexa Tartaglini, Peter Hase +2
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…
Are language models rational? The case of coherence norms and belief revision
Thomas Hofweber, Peter Hase, Elias Stengel-Eskin +1
Do norms of rationality apply to machine learning models, in particular language models? In this paper we investigate this question by focusing on a special subset of rational norm…
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
Vaidehi Patil, Yi-Lin Sung, Peter Hase +3
LLMs trained on massive datasets may inadvertently acquire sensitive information such as personal details and potentially harmful content. This risk is further heightened in multim…
Teaching Models to Balance Resisting and Accepting Persuasion
Elias Stengel-Eskin, Peter Hase, Mohit Bansal
Large language models (LLMs) are susceptible to persuasion, which can pose risks when models are faced with an adversarial interlocutor. We take a first step towards defending mode…
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
Elias Stengel-Eskin, Peter Hase, Mohit Bansal
When answering questions, LLMs can convey not only an answer, but a level of confidence about the answer being correct. This includes explicit confidence markers (e.g. giving a num…
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
Peter Hase, Thomas Hofweber, Xiang Zhou +2
The model editing problem concerns how language models should learn new facts about the world over time. While empirical research on model editing has drawn widespread attention, t…