activity
20172025
most citedScaling Language Models: Methods, Analysis & Insights from Training Gopher

243 citations · 663 across the 15 of their papers we have counts for

collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG202223 cited

Solving math word problems with process- and outcome-based feedback

Jonathan Uesato, Nate Kushman, Ramana Kumar +6

Recent work has shown that asking language models to generate reasoning steps improves performance on many reasoning tasks. When moving beyond prompting, this raises the question o…

cs.LG202216 cited

Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals

Rohin Shah, Vikrant Varma, Ramana Kumar +4

The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming,…

cs.LG2022133 cited

Improving alignment of dialogue agents via targeted human judgements

Amelia Glaese, Nat McAleese, Maja Trębacz +31

We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…

cs.LG2021

An Empirical Investigation of Learning from Biased Toxicity Labels

Neel Nanda, Jonathan Uesato, Sven Gowal

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only…

cs.LG20202 cited

Avoiding Tampering Incentives in Deep RL via Decoupled Approval

Jonathan Uesato, Ramana Kumar, Victoria Krakovna +3

How can we design agents that pursue a given objective when all feedback mechanisms are influenceable by the agent? Standard RL algorithms assume a secure reward function, and can…

cs.LG20203 cited

REALab: An Embedded Perspective on Tampering

Ramana Kumar, Jonathan Uesato, Richard Ngo +3

This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL). REALab is designed to model the structure of tampering problems that may arise…