51 citations · 90 across the 2 of their papers we have counts for
1 paper · 1 filter
Wesley Chung, Valentin Thomas, Marlos C. Machado +1
Bandit and reinforcement learning (RL) problems can often be framed as optimization problems where the goal is to maximize average performance while having access only to stochasti…