23 citations · 55 across the 7 of their papers we have counts for
5 papers · 1 filter
Solving math word problems with process- and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar +6
Recent work has shown that asking language models to generate reasoning steps improves performance on many reasoning tasks. When moving beyond prompting, this raises the question o…
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Rohin Shah, Vikrant Varma, Ramana Kumar +4
The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming,…
Safe Deep RL in 3D Environments using Human Feedback
Matthew Rahtz, Vikrant Varma, Ramana Kumar +3
Agents should avoid unsafe behaviour during both training and deployment. This typically requires a simulator and a procedural specification of unsafe behaviour. Unfortunately, a s…
Avoiding Tampering Incentives in Deep RL via Decoupled Approval
Jonathan Uesato, Ramana Kumar, Victoria Krakovna +3
How can we design agents that pursue a given objective when all feedback mechanisms are influenceable by the agent? Standard RL algorithms assume a secure reward function, and can…
REALab: An Embedded Perspective on Tampering
Ramana Kumar, Jonathan Uesato, Richard Ngo +3
This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL). REALab is designed to model the structure of tampering problems that may arise…