1 paper
Anish Mudide, Joshua Engels, Eric J. Michaud +2
Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features…