2 papers
cs.LG2026
The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2
Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating…
cs.LG2024
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
Jatin Nainani, Sankaran Vaidyanathan, AJ Yeung +2
Mechanistic interpretability aims to understand the inner workings of large neural networks by identifying circuits, or minimal subgraphs within the model that implement algorithms…