1 paper
Piotr Jedryszek, Oliver M. Crook
Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random…