11 citations · 16 across the 3 of their papers we have counts for
3 papers
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…
Learning Universal Predictors
Jordi Grau-Moya, Tim Genewein, Marcus Hutter +8
Meta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data. Broad exposure to different tasks leads to versatile represe…
Learning to Deceive in Multi-Agent Hidden Role Games
Matthew Aitchison, Lyndon Benke, Penny Sweetser
Deception is prevalent in human social settings. However, studies into the effect of deception on reinforcement learning algorithms have been limited to simplistic settings, restri…