17 citations · 21 across the 4 of their papers we have counts for
1 paper · 1 filter
S. A. Murphy, Y. Deng, E. B. Laber +3
We develop an off-policy actor-critic algorithm for learning an optimal policy from a training set composed of data from multiple individuals. This algorithm is developed with a vi…