1 paper · 1 filter
Shun Shao, Binxu Wang, Shay B. Cohen +2
Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expensive, model-specific, and dif…