17 citations · 17 across the 3 of their papers we have counts for
1 paper · 1 filter
Leo Gao, Achyuta Rajaram, Jacob Coxon +3
Finding human-understandable circuits in language models is a central goal of the field of mechanistic interpretability. We train models to have more understandable circuits by con…