145 citations · 176 across the 8 of their papers we have counts for
5 papers · 1 filter
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…
The Hydra Effect: Emergent Self-repair in Language Model Computations
Thomas McGrath, Matthew Rahtz, Janos Kramar +2
We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one att…
Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
Tom Lieberum, Matthew Rahtz, János Kramár +4
\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the stat…
Safe Deep RL in 3D Environments using Human Feedback
Matthew Rahtz, Vikrant Varma, Ramana Kumar +3
Agents should avoid unsafe behaviour during both training and deployment. This typically requires a simulator and a procedural specification of unsafe behaviour. Unfortunately, a s…
An Extensible Interactive Interface for Agent Design
Matthew Rahtz, James Fang, Anca D. Dragan +1
In artificial intelligence, we often specify tasks through a reward function. While this works well in some settings, many tasks are hard to specify this way. In deep reinforcement…