1 paper
Jonathn Chang, Arya Datla, Ziv Goldfeld
Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with the low-level computation realized by a neural network th…