activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

Qinan Yu, Alexa Tartaglini, Peter Hase +2

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…

cs.CL2025

Are language models rational? The case of coherence norms and belief revision

Thomas Hofweber, Peter Hase, Elias Stengel-Eskin +1

Do norms of rationality apply to machine learning models, in particular language models? In this paper we investigate this question by focusing on a special subset of rational norm…

cs.CL2025

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation

Vaidehi Patil, Yi-Lin Sung, Peter Hase +3

LLMs trained on massive datasets may inadvertently acquire sensitive information such as personal details and potentially harmful content. This risk is further heightened in multim…

cs.CL2025

Teaching Models to Balance Resisting and Accepting Persuasion

Elias Stengel-Eskin, Peter Hase, Mohit Bansal

Large language models (LLMs) are susceptible to persuasion, which can pose risks when models are faced with an adversarial interlocutor. We take a first step towards defending mode…

cs.CL2024

LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models

Elias Stengel-Eskin, Peter Hase, Mohit Bansal

When answering questions, LLMs can convey not only an answer, but a level of confidence about the answer being correct. This includes explicit confidence markers (e.g. giving a num…

cs.CL2024

Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?

Peter Hase, Thomas Hofweber, Xiang Zhou +2

The model editing problem concerns how language models should learn new facts about the world over time. While empirical research on model editing has drawn widespread attention, t…