1 paper
Sean P. Fillingham, Andrew Gordon, Peter Lai +3
Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs…