34 citations · 53 across the 4 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.LG2023★ 3 cited
The Hydra Effect: Emergent Self-repair in Language Model Computations
Thomas McGrath, Matthew Rahtz, Janos Kramar +2
We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one att…
cs.LG2023
Tracr: Compiled Transformers as a Laboratory for Interpretability
David Lindner, János Kramár, Sebastian Farquhar +3
We show how to "compile" human-readable programs into standard decoder-only transformer models. Our compiler, Tracr, generates models with known structure. This structure can be us…