1 paper
Min Zeng, Yichen Zhang, Xiaofeng Shao
Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a step…