1 paper · 1 filter
Daniel Ezer, Alon Peled-Cohen, Yishay Mansour
We study the stochastic linear bandits with parameter noise model, in which the reward of action a is a⊤I^¸ where I^¸ is sampled i.i.d. We show a regret upper bound of $\w…