1 paper · 1 filter
Hengtong Lu, Victor Shea-Jay Huang, Chengmin Yang +4
Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated maneuver and the policy cannot r…