43 citations · 61 across the 3 of their papers we have counts for
3 papers
cs.LG2022★ 16 cited
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Rohin Shah, Vikrant Varma, Ramana Kumar +4
The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming,…
cs.LG2022★ 2 cited
Safe Deep RL in 3D Environments using Human Feedback
Matthew Rahtz, Vikrant Varma, Ramana Kumar +3
Agents should avoid unsafe behaviour during both training and deployment. This typically requires a simulator and a procedural specification of unsafe behaviour. Unfortunately, a s…
cs.LG2021★ 43 cited
Imitating Interactive Intelligence
Josh Abramson, Arun Ahuja, Iain Barr +26
A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through…