collaborators

11 papers

cs.CL2026

Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility

Shramay Palta, Peter Rankel, Sarah Wiegreffe +1

We investigate the degree to which human (and LLM) plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments fo…

cs.SD2026

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3

Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability…

cs.CL2026

Localizing Anchoring Pathways in Language Models

Hillary N. Owusu, Sarah Wiegreffe, Naomi H. Feldman

Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside…

cs.LG2026

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe +1

Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making rel…

cs.LG2026

Interpretability Can Be Actionable

Hadas Orgad, Fazl Barez, Tal Haklay +9

Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…

cs.LG2026

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha

Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works-- speci…