1 paper · 1 filter
Chenlu Ye, Yujia Jin, Alekh Agarwal +1
Typical contextual bandit algorithms assume that the rewards at each round lie in some fixed range [0,R], and their regret scales polynomially with this reward range R. Howeve…