1 paper
Taehyun Hwang, Kyuwook Chai, Min-hwan Oh
We consider a contextual combinatorial bandit problem where in each round a learning agent selects a subset of arms and receives feedback on the selected arms according to their sc…