48 citations · 88 across the 3 of their papers we have counts for
3 papers
Engineering Monosemanticity in Toy Models
Adam S. Jermyn, Nicholas Schiefer, Evan Hubinger
In some neural networks, individual neurons correspond to natural ``features'' in the input. Such \emph{monosemantic} neurons are of great help in interpretability studies, as they…
Measuring Progress on Scalable Oversight for Large Language Models
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez +43
Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on m…
Toy Models of Superposition
Nelson Elhage, Tristan Hume, Catherine Olsson +13
Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This…