1 paper
Ayush Sekhari, Karthik Sridharan, Wen Sun +1
We consider the problem of contextual bandits and imitation learning, where the learner lacks direct knowledge of the executed action's reward. Instead, the learner can actively qu…