25 citations · 40 across the 3 of their papers we have counts for
5 papers
Simple Agent, Complex Environment: Efficient Reinforcement Learning with Agent States
Shi Dong, Benjamin Van Roy, Zhengyuan Zhou
We design a simple reinforcement learning (RL) agent that implements an optimistic version of -learning and establish through regret analysis that this agent can operate with so…
Provably Efficient Reinforcement Learning with Aggregated States
Shi Dong, Benjamin Van Roy, Zhengyuan Zhou
We establish that an optimistic variant of Q-learning applied to a fixed-horizon episodic Markov decision process with an aggregated state representation incurs regret $\tilde{\mat…
Comments on the Du-Kakade-Wang-Yang Lower Bounds
Benjamin Van Roy, Shi Dong
Du, Kakade, Wang, and Yang recently established intriguing lower bounds on sample complexity, which suggest that reinforcement learning with a misspecified representation is intrac…
On the Performance of Thompson Sampling on Logistic Bandits
Shi Dong, Tengyu Ma, Benjamin Van Roy
We study the logistic bandit, in which rewards are binary with success probability and actions and coefficients are within the …
An Information-Theoretic Analysis for Thompson Sampling with Many Actions
Shi Dong, Benjamin Van Roy
Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty. However, this dependence is through entropy, which can beco…