8 citations
3 papers
stat.ML2026★ 6 cited
A single algorithm for both restless and rested rotting bandits
Julien Seznec, Pierre Ménard, Alessandro Lazaric +1
In many application domains (e.g., recommender systems, intelligent tutoring systems), the rewards associated to the actions tend to decrease over time. This decay is either caused…
stat.ML2026★ 1 cited
Adaptive multi-fidelity optimization with fast learning rates
Come Fiegel, Victor Gabillon, Michal Valko
In multi-fidelity optimization, biased approximations of varying costs of the target function are available. This paper studies the problem of optimizing a locally smooth function…
cs.LG2026★ 8 cited
Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
Jean Tarbouriech, Matteo Pirotta, Michal Valko +1
We study the sample complexity of learning an -optimal policy in the Stochastic Shortest Path (SSP) problem. We first derive sample complexity bounds when the learner has acces…