492 citations · 566 across the 3 of their papers we have counts for
1 paper · 1 filter
Sid Black, Lee Sharkey, Leo Grinsztajn +8
Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Pre…