1 paper
J. G. Dai, Mark Gluzman
The policy improvement bound on the difference of the discounted returns plays a crucial role in the theoretical justification of the trust-region policy optimization (TRPO) algori…