1 paper
Qin Ding, Cho-Jui Hsieh, James Sharpnack
Classic contextual bandit algorithms for linear models, such as LinUCB, assume that the reward distribution for an arm is modeled by a stationary linear regression. When the linear…