112 citations · 112 across the 4 of their papers we have counts for
1 paper · 2 filters
Ashim Dhor, Pin-Yu Chen
Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method t…