activity
20182026
most citedMIB: A Mechanistic Interpretability Benchmark

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL2026

Localizing Anchoring Pathways in Language Models

Hillary N. Owusu, Sarah Wiegreffe, Naomi H. Feldman

Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside…

cs.CL2026

Quantifying the Gap between Understanding and Generation within Unified Multimodal Models

Chenlong Wang, Yuhang Chen, Zhihan Hu +4

Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are gen…

cs.CL2025

Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility

Shramay Palta, Peter Rankel, Sarah Wiegreffe +1

We investigate the degree to which human (and LLM) plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments fo…

cs.CL2025

On Linear Representations and Pretraining Data Frequency in Language Models

Jack Merullo, Noah A. Smith, Sarah Wiegreffe +1

Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work f…

cs.CL2024

Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning

Shramay Palta, Nishant Balepur, Peter Rankel +3

Questions involving commonsense reasoning about everyday situations often admit many or answers. In contrast, multiple-choice question (MCQ…

cs.CL20241 cited

Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions

Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov +2

Multiple-choice question answering (MCQA) is a key competence of performant transformer language models that is tested by mainstream benchmarks. However, recent evidence shows that…