538 citations · 1.8k across the 57 of their papers we have counts for
4 papers · 1 filter
Policy evaluation from a single path: Multi-step methods, mixing and mis-specification
Yaqi Duan, Martin J. Wainwright
We study non-parametric estimation of the value function of an infinite-horizon -discounted Markov reward process (MRP) using observations from a single trajectory. We provide n…
Off-policy estimation of linear functionals: Non-asymptotic theory for semi-parametric efficiency
Wenlong Mou, Martin J. Wainwright, Peter L. Bartlett
The problem of estimating a linear functional based on observational data is canonical in both the causal inference and bandit literatures. We analyze a broad class of two-stage pr…
A new similarity measure for covariate shift with applications to nonparametric regression
Reese Pathak, Cong Ma, Martin J. Wainwright
We study covariate shift in the context of nonparametric regression. We introduce a new measure of distribution mismatch between the source and target distributions that is based o…
Instance-Dependent Confidence and Early Stopping for Reinforcement Learning
Koulik Khamaru, Eric Xia, Martin J. Wainwright +1
Various algorithms for reinforcement learning (RL) exhibit dramatic variation in their convergence rates as a function of problem structure. Such problem-dependent behavior is not…