1 paper
Eslam Zaher, Maciej Trzaskowski, Quan Nguyen +1
Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of interpretable dictionary atoms,…