1 paper
David Chanin, Tomáš Dulka, Adrià Garriga-Alonso
It is assumed that sparse autoencoders (SAEs) decompose polysemantic activations into interpretable linear directions, as long as the activations are composed of sparse linear comb…