3 papers
cs.LG2026
Scale-free adaptive planning for deterministic dynamics & discounted rewards
Peter L. Bartlett, Victor Gabillon, Jennifer Healey +1
We address the problem of planning in an environment with deterministic dynamics and stochastic rewards with discounted returns. The optimal value function is not known, nor are th…
stat.ML2026
Adaptive multi-fidelity optimization with fast learning rates
Come Fiegel, Victor Gabillon, Michal Valko
In multi-fidelity optimization, biased approximations of varying costs of the target function are available. This paper studies the problem of optimizing a locally smooth function…
stat.ML2026
Best of both worlds: Stochastic & adversarial best-arm identification
Yasin Abbasi-Yadkori, Peter L. Bartlett, Victor Gabillon +2
We study bandit best-arm identification with arbitrary and potentially adversarial rewards. A simple random uniform learner obtains the optimal rate of error in the adversarial sce…