most citedDiscovering Agents

7 citations · 7 across the 1 of their papers we have counts for

collaborators

5 papers

cs.LG20246 cited

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton, Noah Y. Siegel, János Kramár +8

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, whe…

cs.CY202451 cited

The Ethics of Advanced AI Assistants

Iason Gabriel, Arianna Manzini, Geoff Keeling +54

This paper focuses on the opportunities and the ethical and societal risks posed by advanced AI assistants. We define advanced AI assistants as artificial agents with natural langu…

cs.CY20249 cited

A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI

Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery +17

Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generat…

cs.LG20232 cited

Explaining grokking through circuit efficiency

Vikrant Varma, Rohin Shah, Zachary Kenton +2

One of the most surprising puzzles in neural network generalisation is grokking: a network with perfect training accuracy but poor generalisation will, upon further training, trans…

cs.AI20227 cited

Discovering Agents

Zachary Kenton, Ramana Kumar, Sebastian Farquhar +3

Causal models of agents have been used to analyse the safety aspects of machine learning systems. But identifying agents is non-trivial -- often the causal model is just assumed by…