5 citations · 5 across the 2 of their papers we have counts for
1 paper · 1 filter
Bart Bussmann, Patrick Leask, Neel Nanda
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting language model activations by decomposing them into sparse, interpretable features. A popular approach i…