1 paper · 1 filter
Cristina P. Martin-Linares, Jonathan P. Ling
Sparse autoencoders (SAEs) aim to disentangle model activations into monosemantic, human-interpretable features. In practice, learned features are often redundant and vary across t…