1 paper
Nicolas Gutowski, Fabien Chhel, Alexandre Letard +1
We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of k arms and observes their d-dimensional reward vectors under semi-ban…