most citedDoes Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

6 citations · 15 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG20242 cited

Improving Dictionary Learning with Gated Sparse Autoencoders

Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith +5

Recent work has found that sparse autoencoders (SAEs) are an effective technique for unsupervised discovery of interpretable features in language models' (LMs) activations, by find…

cs.LG20232 cited

Explaining grokking through circuit efficiency

Vikrant Varma, Rohin Shah, Zachary Kenton +2

One of the most surprising puzzles in neural network generalisation is grokking: a network with perfect training accuracy but poor generalisation will, upon further training, trans…

cs.LG20233 cited

The Hydra Effect: Emergent Self-repair in Language Model Computations

Thomas McGrath, Matthew Rahtz, Janos Kramar +2

We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one att…

cs.LG20236 cited

Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Tom Lieberum, Matthew Rahtz, János Kramár +4

\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the stat…

cs.AI20232 cited

Power-seeking can be probable and predictive for trained agents

Victoria Krakovna, Janos Kramar

Power-seeking behavior is a key source of risk from advanced AI, but our theoretical understanding of this phenomenon is relatively limited. Building on existing theoretical result…