7 citations · 7 across the 1 of their papers we have counts for
5 papers
On scalable oversight with weak LLMs judging strong LLMs
Zachary Kenton, Noah Y. Siegel, János Kramár +8
Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, whe…
The Ethics of Advanced AI Assistants
Iason Gabriel, Arianna Manzini, Geoff Keeling +54
This paper focuses on the opportunities and the ethical and societal risks posed by advanced AI assistants. We define advanced AI assistants as artificial agents with natural langu…
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery +17
Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generat…
Explaining grokking through circuit efficiency
Vikrant Varma, Rohin Shah, Zachary Kenton +2
One of the most surprising puzzles in neural network generalisation is grokking: a network with perfect training accuracy but poor generalisation will, upon further training, trans…
Discovering Agents
Zachary Kenton, Ramana Kumar, Sebastian Farquhar +3
Causal models of agents have been used to analyse the safety aspects of machine learning systems. But identifying agents is non-trivial -- often the causal model is just assumed by…