1 paper
Omer Amichay, Yishay Mansour
We consider non-stationary multi-arm bandit (MAB) where the expected reward of each action follows a linear function of the number of times we executed the action. Our main result…