1 paper · 1 filter
Leo Gao, Achyuta Rajaram, Jacob Coxon +3
Finding human-understandable circuits in language models is a central goal of the field of mechanistic interpretability. We train models to have more understandable circuits by con…