39 citations · 47 across the 4 of their papers we have counts for
6 papers
Learning to Stop with Surprisingly Few Samples
Daniel Russo, Assaf Zeevi, Tianyi Zhang
We consider a discounted infinite horizon optimal stopping problem. If the underlying distribution is known a priori, the solution of this problem is obtained via dynamic programmi…
Policy Gradient Optimization of Thompson Sampling Policies
Seungki Min, Ciamac C. Moallemi, Daniel J. Russo
We study the use of policy gradient algorithms to optimize over a class of generalized Thompson sampling policies. Our central insight is to view the posterior parameter sampled by…
A Note on the Equivalence of Upper Confidence Bounds and Gittins Indices for Patient Agents
Daniel Russo
This note gives a short, self-contained, proof of a sharp connection between Gittins indices and Bayesian upper confidence bound algorithms. I consider a Gaussian multi-armed bandi…
A Finite Time Analysis of Temporal Difference Learning With Linear Function Approximation
Jalaj Bhandari, Daniel Russo, Raghav Singal
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is o…
Satisficing in Time-Sensitive Bandit Learning
Daniel Russo, Benjamin Van Roy
Much of the recent literature on bandit learning focuses on algorithms that aim to converge on an optimal action. One shortcoming is that this orientation does not account for time…
Improving the Expected Improvement Algorithm
Chao Qin, Diego Klabjan, Daniel Russo
The expected improvement (EI) algorithm is a popular strategy for information collection in optimization under uncertainty. The algorithm is widely known to be too greedy, but neve…