1 paper · 1 filter
Pierre Boudart, Pierre Gaillard, Alessandro Rudi
We consider the multinomial logistic bandit problem in which a learner interacts with an environment by selecting actions to maximize expected rewards based on probabilistic feedba…