51 citations · 51 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 51 cited
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Hoagy Cunningham, Aidan Ewart, Logan Riggs +2
One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct conte…
cs.LG2023
A technical note on bilinear layers for interpretability
Lee Sharkey
The ability of neural networks to represent more features than neurons makes interpreting them challenging. This phenomenon, known as superposition, has spurred efforts to find arc…