activity
20192026
most citedGemma 2: Improving Open Language Models at a Practical Size

145 citations · 176 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG202411 cited

Evaluating Frontier Models for Dangerous Capabilities

Mary Phuong, Matthew Aitchison, Elliot Catt +24

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…

cs.LG20233 cited

The Hydra Effect: Emergent Self-repair in Language Model Computations

Thomas McGrath, Matthew Rahtz, Janos Kramar +2

We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one att…

cs.LG20236 cited

Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Tom Lieberum, Matthew Rahtz, János Kramár +4

\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the stat…

cs.LG20222 cited

Safe Deep RL in 3D Environments using Human Feedback

Matthew Rahtz, Vikrant Varma, Ramana Kumar +3

Agents should avoid unsafe behaviour during both training and deployment. This typically requires a simulator and a procedural specification of unsafe behaviour. Unfortunately, a s…

cs.LG2019

An Extensible Interactive Interface for Agent Design

Matthew Rahtz, James Fang, Anca D. Dragan +1

In artificial intelligence, we often specify tasks through a reward function. While this works well in some settings, many tasks are hard to specify this way. In deep reinforcement…