23 citations · 23 across the 1 of their papers we have counts for
1 paper
Lauro Langosco, Jack Koch, Lee Sharkey +3
We study goal misgeneralization, a type of out-of-distribution generalization failure in reinforcement learning (RL). Goal misgeneralization failures occur when an RL agent retains…