107 citations · 197 across the 7 of their papers we have counts for
1 paper · 2 filters
Pablo Samuel Castro, Shijian Li, Daqing Zhang
We consider the problem of learning to behave optimally in a Markov Decision Process when a reward function is not specified, but instead we have access to a set of demonstrators o…