91 citations · 167 across the 11 of their papers we have counts for
13 papers
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Rohin Shah, Vikrant Varma, Ramana Kumar +4
The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming,…
An Empirical Investigation of Representation Learning for Imitation
Xin Chen, Sam Toyer, Cody Wild +9
Imitation learning often needs a large demonstration set in order to handle the full range of situations that an agent might find itself in during deployment. However, collecting e…
Retrospective on the 2021 BASALT Competition on Learning from Human Feedback
Rohin Shah, Steven H. Wang, Cody Wild +13
We held the first-ever MineRL Benchmark for Agents that Solve Almost-Lifelike Tasks (MineRL BASALT) Competition at the Thirty-fifth Conference on Neural Information Processing Syst…
The MineRL BASALT Competition on Learning from Human Feedback
Rohin Shah, Cody Wild, Steven H. Wang +10
The last decade has seen a significant increase of interest in deep learning research, with many public successes that have demonstrated its potential. As such, these systems are n…
Learning What To Do by Simulating the Past
David Lindner, Rohin Shah, Pieter Abbeel +1
Since reward functions are hard to specify, recent work has focused on learning policies from human feedback. However, such approaches are impeded by the expense of acquiring such…
Combining Reward Information from Multiple Sources
Dmitrii Krasheninnikov, Rohin Shah, Herke van Hoof
Given two sources of evidence about a latent variable, one can combine the information from both by multiplying the likelihoods of each piece of evidence. However, when one or both…