2 papers
stat.ML2026
Replicability is Asymptotically Free in Multi-armed Bandits
Junpei Komiyama, Shinji Ito, Yuichi Yoshida +1
We consider a replicable stochastic multi-armed bandit algorithm that ensures, with high probability, that the algorithm's sequence of actions is not affected by the randomness inh…
cs.LG2025
Data-dependent Bounds with -Optimal Best-of-Both-Worlds Guarantees in Multi-Armed Bandits using Stability-Penalty Matching
Quan Nguyen, Shinji Ito, Junpei Komiyama +1
Existing data-dependent and best-of-both-worlds regret bounds for multi-armed bandits problems have limited adaptivity as they are either data-dependent but not best-of-both-worlds…