1 paper
Maheep Chaudhary, Atticus Geiger
A popular new method in mechanistic interpretability is to train high-dimensional sparse autoencoders (SAEs) on neuron activations and use SAE features as the atomic units of analy…