3 papers
stat.ML2025
Tractable Instances of Bilinear Maximization: Implementing LinUCB on Ellipsoids
Raymond Zhang, Hédi Hadiji, Richard Combes
We consider the maximization of over , with convex and an ellipsoid. This…
stat.ML2025
Linear Bandits on Ellipsoids: Minimax Optimal Algorithms
Raymond Zhang, Hedi Hadiji, Richard Combes
We consider linear stochastic bandits where the set of actions is an ellipsoid. We provide the first known minimax optimal algorithm for this problem. We first derive a novel infor…
stat.ML2024
Thompson Sampling For Combinatorial Bandits: Polynomial Regret and Mismatched Sampling Paradox
Raymond Zhang, Richard Combes
We consider Thompson Sampling (TS) for linear combinatorial semi-bandits and subgaussian rewards. We propose the first known TS whose finite-time regret does not scale exponentiall…