3 citations · 3 across the 1 of their papers we have counts for
1 paper
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat +4
Sparse autoencoders (SAEs) are a promising unsupervised approach for identifying causally relevant and interpretable linear features in a language model's (LM) activations. To be u…