1 paper
Moise Blanchard, Steve Hanneke, Patrick Jaillet
We study the fundamental limits of learning in contextual bandits, where a learner's rewards depend on their actions and a known context, which extends the canonical multi-armed ba…