3 citations · 3 across the 1 of their papers we have counts for
1 paper · 1 filter
Michael James McDonald, Dylan Hadfield-Menell
While modern policy optimization methods can do complex manipulation from sensory data, they struggle on problems with extended time horizons and multiple sub-goals. On the other h…