1 paper · 1 filter
Ruben Fernandez-Boullon, Pablo Magariños-Docampo, Javier Perez-Robles
Sparse autoencoders (SAEs) have become central to mechanistic interpretability, decomposing transformer activations into monosemantic features. Yet existing analyses characterise f…