51 citations · 51 across the 3 of their papers we have counts for
1 paper · 2 filters
Patrick Leask, Bart Bussmann, Michael Pearce +5
A common goal of mechanistic interpretability is to decompose the activations of neural networks into features: interpretable properties of the input computed by the model. Sparse…