1 paper
Xiaoyu Zhang, Matthew Chang, Pranav Kumar +1
A common failure mode for policies trained with imitation is compounding execution errors at test time. When the learned policy encounters states that are not present in the expert…