15 citations · 47 across the 6 of their papers we have counts for
Showing stat.MLShow all
2 papers · 1 filter
stat.ML2022★ 1 cited
Policy evaluation from a single path: Multi-step methods, mixing and mis-specification
Yaqi Duan, Martin J. Wainwright
We study non-parametric estimation of the value function of an infinite-horizon -discounted Markov reward process (MRP) using observations from a single trajectory. We provide n…
stat.ML2021★ 4 cited
Optimal policy evaluation using kernel-based temporal difference methods
Yaqi Duan, Mengdi Wang, Martin J. Wainwright
We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized…