11 papers
Unsupervised Causal Abstractions Discovery
Théo Saulus, Simon Lacoste-Julien, Dhanya Sridhar
Causal abstractions formalize when a high-level structural causal model (SCM) captures the interventional behavior of a lower-level SCM. Existing applications of this notion largel…
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Aaron Mueller, Andrew Lee, Shruti Joshi +3
A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically ev…
The Role of Causal Features in Strategic Classification for Robustness and Alignment
Antonio Gois, Sophia Gunluk, Nir Rosenfeld +3
In strategic classification, an institution (e.g., a bank) anticipates adaptation from users who change their features to increase utility in a classification task (e.g., loan repa…
Causality is Key for Interpretability Claims to Generalise
Shruti Joshi, Aaron Mueller, David Klindt +3
Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persist: findings that do not generalise, and…
Demystifying amortized causal discovery with transformers
Francesco Montagna, Max Cairney-Leeming, Dhanya Sridhar +1
Supervised learning for causal discovery from observational data often achieves competitive performance despite seemingly avoiding the explicit assumptions that traditional methods…
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
Shruti Joshi, Théo Saulus, Wieland Brendel +3
Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with known ground-truth factors. These metrics…