Publications (49)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
Raphaël Sarfati, Eric Bigelow, Daniel Wurgaft +6
Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representa…
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
Amir Zur, Atticus Geiger, Ekdeep Singh Lubana +1
When a language model generates text, the selection of individual tokens might lead it down very different reasoning paths, making uncertainty difficult to quantify. In this work,…
Causal Distillation for Language Models
Zhengxuan Wu, Atticus Geiger, Josh Rozner +5
Distillation efforts have led to language models that are more compact and efficient without serious drops in performance. The standard approach to distillation trains a student mo…
CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior
Eldar David Abraham, Karel D'Oosterlinck, Amir Feder +5
The increasing size and complexity of modern ML systems has improved their predictive capabilities but made their behavior harder to explain. Many techniques for model explanation…
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Atticus Geiger, Duligur Ibeling, Amir Zur +8
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications…
Inducing Causal Structure for Interpretable Neural Networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu +5
In many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashio…