12 citations · 19 across the 5 of their papers we have counts for
4 papers · 1 filter
Gaussian Imagination in Bandit Learning
Yueyang Liu, Adithya M. Devraj, Benjamin Van Roy +1
Assuming distributions are Gaussian often facilitates computations that are otherwise intractable. We study the performance of an agent that attains a bounded information ratio wit…
A Bit Better? Quantifying Information for Bandit Learning
Adithya M. Devraj, Benjamin Van Roy, Kuang Xu
The information ratio offers an approach to assessing the efficacy with which an agent balances between exploration and exploitation. Originally, this was defined to be the ratio b…
Q-learning with Uniformly Bounded Variance: Large Discounting is Not a Barrier to Fast Learning
Adithya M. Devraj, Sean P. Meyn
Sample complexity bounds are a common performance metric in the Reinforcement Learning literature. In the discounted cost, infinite horizon setting, all of the known bounds have a…
Zap Q-Learning With Nonlinear Function Approximation
Shuhang Chen, Adithya M. Devraj, Fan Lu +2
Zap Q-learning is a recent class of reinforcement learning algorithms, motivated primarily as a means to accelerate convergence. Stability theory has been absent outside of two res…