1 paper · 1 filter
Piotr Jedryszek, Oliver M. Crook
Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random…