activity
20182021
most citedComments on the Du-Kakade-Wang-Yang Lower Bounds

25 citations · 40 across the 3 of their papers we have counts for

collaborators

5 papers

cs.LG2021

Simple Agent, Complex Environment: Efficient Reinforcement Learning with Agent States

Shi Dong, Benjamin Van Roy, Zhengyuan Zhou

We design a simple reinforcement learning (RL) agent that implements an optimistic version of -learning and establish through regret analysis that this agent can operate with so…

stat.ML201911 cited

Provably Efficient Reinforcement Learning with Aggregated States

Shi Dong, Benjamin Van Roy, Zhengyuan Zhou

We establish that an optimistic variant of Q-learning applied to a fixed-horizon episodic Markov decision process with an aggregated state representation incurs regret $\tilde{\mat…

cs.LG201925 cited

Comments on the Du-Kakade-Wang-Yang Lower Bounds

Benjamin Van Roy, Shi Dong

Du, Kakade, Wang, and Yang recently established intriguing lower bounds on sample complexity, which suggest that reinforcement learning with a misspecified representation is intrac…

stat.ML20194 cited

On the Performance of Thompson Sampling on Logistic Bandits

Shi Dong, Tengyu Ma, Benjamin Van Roy

We study the logistic bandit, in which rewards are binary with success probability and actions and coefficients are within the

stat.ML2018

An Information-Theoretic Analysis for Thompson Sampling with Many Actions

Shi Dong, Benjamin Van Roy

Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty. However, this dependence is through entropy, which can beco…