1 paper
Apurv Shukla, P. R. Kumar
We consider contextual bandit learning under distribution shift when reward vectors are ordered according to a given preference cone. We propose an adaptive-discretization and opti…