12 citations · 23 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 5 cited
Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Neel Nanda, Andrew Lee, Martin Wattenberg
How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023…
cs.LG2023★ 6 cited
Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
Tom Lieberum, Matthew Rahtz, János Kramár +4
\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the stat…
cs.LG2023★ 12 cited
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
Bilal Chughtai, Lawrence Chan, Neel Nanda
Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks. In this work, we study…