A Second-Order Method for Stochastic Bandit Convex Optimisation
arXiv:2302.05371
Abstract
We introduce a simple and efficient algorithm for unconstrained zeroth-order stochastic convex bandits and prove its regret is at most where is the horizon, the dimension and is the radius of a known ball containing the minimiser of the loss.
27 pages