5 citations · 5 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
A technical note on bilinear layers for interpretability
Lee Sharkey
The ability of neural networks to represent more features than neurons makes interpreting them challenging. This phenomenon, known as superposition, has spurred efforts to find arc…
cs.LG2022★ 5 cited
Interpreting Neural Networks through the Polytope Lens
Sid Black, Lee Sharkey, Leo Grinsztajn +8
Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Pre…