751 citations · 1.2k across the 14 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022★ 48 cited
Toy Models of Superposition
Nelson Elhage, Tristan Hume, Catherine Olsson +13
Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This…
gr-qc2022★ 1 cited
Interpreting a Machine Learning Model for Detecting Gravitational Waves
Mohammadtaher Safarzadeh, Asad Khan, E. A. Huerta +1
We describe a case study of translational research, applying interpretability techniques developed for computer vision to machine learning models used to search for and find gravit…