2 papers
cs.LG2020
Stable and Efficient Policy Evaluation
Daoming Lyu, Bo Liu, Matthieu Geist +3
Policy evaluation algorithms are essential to reinforcement learning due to their ability to predict the performance of a policy. However, there are two long-standing issues lying…
cs.LG2017
OTD: (Near)-Optimal Off-Policy TD Learning
Bo Liu, Daoming Lyu, Wen Dong +1
Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their obj…