4 papers
Interactions Between Crosscoder Features: A Compact Proofs Perspective
Dmitry Manning-Coe, Thomas Read, Anna Soligo +4
Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features. Interactions betwee…
Towards a unified and verified understanding of group-operation networks
Wilson Wu, Louis Jaburi, Jacob Drori +1
A recent line of work in mechanistic interpretability has focused on reverse-engineering the computation performed by neural networks trained on the binary operation of finite grou…
Compact Proofs of Model Performance via Mechanistic Interpretability
Jason Gross, Rajashree Agrawal, Thomas Kwa +5
We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guaran…
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration
Chun Hei Yip, Rajashree Agrawal, Lawrence Chan +1
The goal of mechanistic interpretability is discovering simpler, low-rank algorithms implemented by models. While we can compress activations into features, compressing nonlinear f…