317 citations · 769 across the 6 of their papers we have counts for
8 papers
A Human Rights-Based Approach to Responsible AI
Vinodkumar Prabhakaran, Margaret Mitchell, Timnit Gebru +1
Research on fairness, accountability, transparency and ethics of AI-based interventions in society has gained much-needed momentum in recent years. However it lacks an explicit ali…
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
Power to the People? Opportunities and Challenges for Participatory AI
Abeba Birhane, William Isaac, Vinodkumar Prabhakaran +4
Participatory approaches to artificial intelligence (AI) and machine learning (ML) are gaining momentum: the increased attention comes partly with the view that participation opens…
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…
Alignment of Language Agents
Zachary Kenton, Tom Everitt, Laura Weidinger +3
For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for la…
Modelling Cooperation in Network Games with Spatio-Temporal Complexity
Michiel A. Bakker, Richard Everett, Laura Weidinger +4
The real world is awash with multi-agent problems that require collective action by self-interested agents, from the routing of packets across a computer network to the management…