1 paper
Pierre Boudart, Pierre Gaillard, Alessandro Rudi
We consider the multinomial logistic bandit problem in which a learner interacts with an environment by selecting actions to maximize expected rewards based on probabilistic feedba…