activity
20182022
most citedMinimax-Optimal Off-Policy Evaluation with Linear Function Approximation

15 citations · 47 across the 6 of their papers we have counts for

collaborators

9 papers

stat.ML20221 cited

Policy evaluation from a single path: Multi-step methods, mixing and mis-specification

Yaqi Duan, Martin J. Wainwright

We study non-parametric estimation of the value function of an infinite-horizon -discounted Markov reward process (MRP) using observations from a single trajectory. We provide n…

cs.LG20223 cited

Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with Pessimism

Ming Yin, Yaqi Duan, Mengdi Wang +1

Offline reinforcement learning, which seeks to utilize offline/historical data to optimize sequential decision-making strategies, has gained surging prominence in recent studies. D…

stat.ML20214 cited

Optimal policy evaluation using kernel-based temporal difference methods

Yaqi Duan, Mengdi Wang, Martin J. Wainwright

We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized…

cs.LG202114 cited

Risk Bounds and Rademacher Complexity in Batch Reinforcement Learning

Yaqi Duan, Chi Jin, Zhiyuan Li

This paper considers batch Reinforcement Learning (RL) with general value function approximation. Our study investigates the minimal assumptions to reliably estimate/minimize Bellm…

cs.LG202010 cited

Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient

Botao Hao, Yaqi Duan, Tor Lattimore +2

This paper provides a statistical analysis of high-dimensional batch Reinforcement Learning (RL) using sparse linear function approximation. When there is a large number of candida…

cs.LG202015 cited

Minimax-Optimal Off-Policy Evaluation with Linear Function Approximation

Yaqi Duan, Mengdi Wang

This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cum…