380 citations · 475 across the 3 of their papers we have counts for
3 papers
cs.LG2016★ 380 cited
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala +6
In recent years deep reinforcement learning (RL) systems have attained superhuman performance in a number of challenging task domains. However, a major limitation of such applicati…
cs.LG2014★ 82 cited
Bounded Regret for Finite-Armed Structured Bandits
Tor Lattimore, Remi Munos
We study a new type of K-armed bandit problem where the expected return of one arm may depend on the returns of other arms. We present a new algorithm for this general class of pro…
cs.AI2014★ 13 cited
On Minimax Optimal Offline Policy Evaluation
Lihong Li, Remi Munos, Csaba Szepesvari
This paper studies the off-policy evaluation problem, where one aims to estimate the value of a target policy based on a sample of observations collected by another policy. We firs…