9 citations · 9 across the 1 of their papers we have counts for
2 papers
cs.LG2020★ 9 cited
Regret Balancing for Bandit and RL Model Selection
Yasin Abbasi-Yadkori, Aldo Pacchiano, My Phan
We consider model selection in stochastic bandit and reinforcement learning problems. Given a set of base learning algorithms, an effective model selection strategy adapts to the b…
cs.LG2019
Thompson Sampling with Approximate Inference
My Phan, Yasin Abbasi-Yadkori, Justin Domke
We study the effects of approximate inference on the performance of Thompson sampling in the -armed bandit problems. Thompson sampling is a successful algorithm for online decis…