1 paper
Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2
Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating…