1 paper · 1 filter
Ashim Dhor, Pin-Yu Chen
Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method t…