activity
20162022
most citedScaling Language Models: Methods, Analysis & Insights from Training Gopher

243 citations · 364 across the 5 of their papers we have counts for

collaborators

9 papers

cs.CL202254 cited

Teaching language models to support answers with verified quotes

Jacob Menick, Maja Trebacz, Vladimir Mikulik +8

Recent large language models often answer factual questions correctly. But users can't trust any given claim a model makes without fact-checking, because language models can halluc…

cs.CL20227 cited

Uncertainty Estimation for Language Reward Models

Adam Gleave, Geoffrey Irving

Language models can learn a range of capabilities from unsupervised training on text corpora. However, to solve a particular problem (such as text summarization) it is typically ne…

cs.CL202219 cited

Red Teaming Language Models with Language Models

Ethan Perez, Saffron Huang, Francis Song +6

Language Models (LMs) often cannot be deployed because of their potential to harm users in hard-to-predict ways. Prior work identifies harmful behaviors before deployment by using…

cs.CL2022243 cited

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…

cs.AI202141 cited

Alignment of Language Agents

Zachary Kenton, Tom Everitt, Laura Weidinger +3

For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for la…

cs.CL2019

Fine-Tuning Language Models from Human Preferences

Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu +5

Reward learning enables the application of reinforcement learning (RL) to tasks where reward is defined by human judgment, building a model of reward by asking humans questions. Mo…