1 paper · 1 filter
Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2
Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating…