9 papers · 1 filter
Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility
Shramay Palta, Peter Rankel, Sarah Wiegreffe +1
We investigate the degree to which human (and LLM) plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments fo…
Localizing Anchoring Pathways in Language Models
Hillary N. Owusu, Sarah Wiegreffe, Naomi H. Feldman
Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside…
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
Chenlong Wang, Yuhang Chen, Zhihan Hu +4
Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are gen…
On Linear Representations and Pretraining Data Frequency in Language Models
Jack Merullo, Noah A. Smith, Sarah Wiegreffe +1
Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work f…
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov +2
Multiple-choice question answering (MCQA) is a key competence of performant transformer language models that is tested by mainstream benchmarks. However, recent evidence shows that…
The Art of Saying No: Contextual Noncompliance in Language Models
Faeze Brahman, Sachin Kumar, Vidhisha Balachandran +11
Chat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of "unsafe" queries…