2 papers
cs.LG2024
Compact Proofs of Model Performance via Mechanistic Interpretability
Jason Gross, Rajashree Agrawal, Thomas Kwa +5
We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guaran…
cs.LG2024
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration
Chun Hei Yip, Rajashree Agrawal, Lawrence Chan +1
The goal of mechanistic interpretability is discovering simpler, low-rank algorithms implemented by models. While we can compress activations into features, compressing nonlinear f…