243 citations · 417 across the 3 of their papers we have counts for
5 papers
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…
Alignment of Language Agents
Zachary Kenton, Tom Everitt, Laura Weidinger +3
For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for la…
Modelling Cooperation in Network Games with Spatio-Temporal Complexity
Michiel A. Bakker, Richard Everett, Laura Weidinger +4
The real world is awash with multi-agent problems that require collective action by self-interested agents, from the routing of packets across a computer network to the management…
Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences
Raphael Köster, Kevin R. McKee, Richard Everett +7
Game theoretic views of convention generally rest on notions of common knowledge and hyper-rational models of individual behavior. However, decades of work in behavioral economics…