2 papers
cs.LG2026
Interactions Between Crosscoder Features: A Compact Proofs Perspective
Dmitry Manning-Coe, Thomas Read, Anna Soligo +4
Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features. Interactions betwee…
cs.LG2024
Compact Proofs of Model Performance via Mechanistic Interpretability
Jason Gross, Rajashree Agrawal, Thomas Kwa +5
We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guaran…