1 citations · 1 across the 2 of their papers we have counts for
5 papers
Think Before You Lie: How Reasoning Leads to Honesty
Ann Yuan, Asma Ghandeharioun, Carter Blum +6
While existing evaluations of large language models (LLMs) measure deception rates, the underlying conditions that give rise to deceptive behavior are poorly understood. We investi…
When Can Transformers Count to n?
Gilad Yehudai, Haim Kaplan, Guy Dar +4
Large language models based on the transformer architecture can solve highly complex tasks, yet their fundamental limitations on simple algorithmic problems remain poorly understoo…
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
Carter Blum, Katja Filippova, Ann Yuan +8
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they hallucinate when asked in one language about facts expressed in a different language during traini…
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
Michael A. Lepori, Michael C. Mozer, Asma Ghandeharioun
The profound success of transformer-based language models can largely be attributed to their ability to integrate relevant contextual information from an input sequence in order to…
Towards Unifying Interpretability and Control: Evaluation via Intervention
Usha Bhalla, Suraj Srinivas, Asma Ghandeharioun +1
With the growing complexity and capability of large language models, a need to understand model reasoning has emerged, often motivated by an underlying goal of controlling and alig…