1 paper · 1 filter
Evan Ryan Gunter, Yevgeny Liokumovich, Victoria Krakovna
We investigate the question: if an AI agent is known to be safe in one setting, is it also safe in a new setting similar to the first? This is a core question of AI alignment--we t…