1 paper
Manjesh K. Hanawal, Hao Liu, Henghui Zhu +1
We consider the problem of learning a policy for a Markov decision process consistent with data captured on the state-actions pairs followed by the policy. We assume that the polic…