papers

Publications (49)

cs.CL2026

The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors

Raphaël Sarfati, Eric Bigelow, Daniel Wurgaft +6

Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representa…

cs.CL2025

Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics

Amir Zur, Atticus Geiger, Ekdeep Singh Lubana +1

When a language model generates text, the selection of individual tokens might lead it down very different reasoning paths, making uncertainty difficult to quantify. In this work,…

cs.CL2022

Causal Distillation for Language Models

Zhengxuan Wu, Atticus Geiger, Josh Rozner +5

Distillation efforts have led to language models that are more compact and efficient without serious drops in performance. The standard approach to distillation trains a student mo…

cs.CL2022

CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior

Eldar David Abraham, Karel D'Oosterlinck, Amir Feder +5

The increasing size and complexity of modern ML systems has improved their predictive capabilities but made their behavior harder to explain. Many techniques for model explanation…

cs.AI2025

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Atticus Geiger, Duligur Ibeling, Amir Zur +8

Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications…

cs.LG2022

Inducing Causal Structure for Interpretable Neural Networks

Atticus Geiger, Zhengxuan Wu, Hanson Lu +5

In many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashio…