2 citations · 2 across the 1 of their papers we have counts for
1 paper
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith +5
Recent work has found that sparse autoencoders (SAEs) are an effective technique for unsupervised discovery of interpretable features in language models' (LMs) activations, by find…