1 paper · 1 filter
Michael Li, Nishant Subramani
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by me…