1 paper
Nirandika Wanigasekara, Christina Lee Yu
Consider a nonparametric contextual multi-arm bandit problem where each arm a∈[K] is associated to a nonparametric reward function fa:[0,1]→R mapping from co…