53 citations · 64 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 11 cited
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…
cs.LG2022★ 53 cited
Formal Algorithms for Transformers
Mary Phuong, Marcus Hutter
This document aims to be a self-contained, mathematically precise overview of transformer architectures and algorithms (*not* results). It covers what transformers are, how they ar…