1 paper
Xingguo Chen, Zhaohui Wu, Jinguo Ye +5
Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize samp…