1 citations · 1 across the 10 of their papers we have counts for
3 papers · 1 filter
Language Models Struggle to Use Representations Learned In-Context
Michael A. Lepori, Tal Linzen, Ann Yuan +1
Though large language models (LLMs) have enabled great success across a wide variety of tasks, they still appear to fall short of one of the loftier goals of artificial intelligenc…
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
Carter Blum, Katja Filippova, Ann Yuan +8
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language dur…
Who's asking? User personas and the mechanics of latent misalignment
Asma Ghandeharioun, Ann Yuan, Marius Guerard +3
Despite investments in improving model safety, studies show that misaligned capabilities remain latent in safety-tuned models. In this work, we shed light on the mechanics of this…