1 paper · 2 filters
Arun Verma, Indrajit Saha, Makoto Yokoo +1
This paper considers a contextual bandit problem involving multiple agents, where a learner sequentially observes the contexts and the agent's reported arms, and then selects the a…