25 citations · 40 across the 3 of their papers we have counts for
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2019★ 11 cited
Provably Efficient Reinforcement Learning with Aggregated States
Shi Dong, Benjamin Van Roy, Zhengyuan Zhou
We establish that an optimistic variant of Q-learning applied to a fixed-horizon episodic Markov decision process with an aggregated state representation incurs regret $\tilde{\mat…
stat.ML2019★ 4 cited
On the Performance of Thompson Sampling on Logistic Bandits
Shi Dong, Tengyu Ma, Benjamin Van Roy
We study the logistic bandit, in which rewards are binary with success probability and actions and coefficients are within the …
stat.ML2018
An Information-Theoretic Analysis for Thompson Sampling with Many Actions
Shi Dong, Benjamin Van Roy
Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty. However, this dependence is through entropy, which can beco…