39 citations · 47 across the 4 of their papers we have counts for
Showing 2018Show all
2 papers · 1 filter
cs.LG2018
A Finite Time Analysis of Temporal Difference Learning With Linear Function Approximation
Jalaj Bhandari, Daniel Russo, Raghav Singal
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is o…
cs.LG2018
Satisficing in Time-Sensitive Bandit Learning
Daniel Russo, Benjamin Van Roy
Much of the recent literature on bandit learning focuses on algorithms that aim to converge on an optimal action. One shortcoming is that this orientation does not account for time…