15 citations · 47 across the 6 of their papers we have counts for
9 papers
Policy evaluation from a single path: Multi-step methods, mixing and mis-specification
Yaqi Duan, Martin J. Wainwright
We study non-parametric estimation of the value function of an infinite-horizon -discounted Markov reward process (MRP) using observations from a single trajectory. We provide n…
Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with Pessimism
Ming Yin, Yaqi Duan, Mengdi Wang +1
Offline reinforcement learning, which seeks to utilize offline/historical data to optimize sequential decision-making strategies, has gained surging prominence in recent studies. D…
Optimal policy evaluation using kernel-based temporal difference methods
Yaqi Duan, Mengdi Wang, Martin J. Wainwright
We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized…
Risk Bounds and Rademacher Complexity in Batch Reinforcement Learning
Yaqi Duan, Chi Jin, Zhiyuan Li
This paper considers batch Reinforcement Learning (RL) with general value function approximation. Our study investigates the minimal assumptions to reliably estimate/minimize Bellm…
Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient
Botao Hao, Yaqi Duan, Tor Lattimore +2
This paper provides a statistical analysis of high-dimensional batch Reinforcement Learning (RL) using sparse linear function approximation. When there is a large number of candida…
Minimax-Optimal Off-Policy Evaluation with Linear Function Approximation
Yaqi Duan, Mengdi Wang
This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cum…