17 citations · 17 across the 1 of their papers we have counts for
1 paper · 1 filter
Chenlu Ye, Yujia Jin, Alekh Agarwal +1
Typical contextual bandit algorithms assume that the rewards at each round lie in some fixed range [0,R], and their regret scales polynomially with this reward range R. Howeve…