11 papers
Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility
Shramay Palta, Peter Rankel, Sarah Wiegreffe +1
We investigate the degree to which human (and LLM) plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments fo…
A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3
Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability…
Localizing Anchoring Pathways in Language Models
Hillary N. Owusu, Sarah Wiegreffe, Naomi H. Feldman
Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside…
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe +1
Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making rel…
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha
Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works-- speci…