6 citations · 9 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 3 cited
The Hydra Effect: Emergent Self-repair in Language Model Computations
Thomas McGrath, Matthew Rahtz, Janos Kramar +2
We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one att…
cs.LG2023★ 6 cited
Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
Tom Lieberum, Matthew Rahtz, János Kramár +4
\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the stat…