An Information-Theoretic Approach to Minimax Regret in Partial Monitoring
arXiv:1902.00470
Abstract
We prove a new minimax theorem connecting the worst-case Bayesian regret and minimax regret under partial monitoring with no assumptions on the space of signals or decisions of the adversary. We then generalise the information-theoretic tools of Russo and Van Roy (2016) for proving Bayesian regret bounds and combine them with the minimax theorem to derive minimax regret bounds for various partial monitoring settings. The highlight is a clean analysis of `non-degenerate easy' and `hard' finite partial monitoring, with new regret bounds that are independent of arbitrarily large game-dependent constants. The power of the generalised machinery is further demonstrated by proving that the minimax regret for k-armed adversarial bandits is at most sqrt{2kn}, improving on existing results by a factor of 2. Finally, we provide a simple analysis of the cops and robbers game, also improving best known constants.
29 pages, to appear in COLT 2019
References in corpus (3)
Cited by in corpus (9)
- Connections Between Mirror Descent, Thompson Sampling and the Information Ratio
- Information Directed Sampling for Linear Partial Monitoring
- Reinforcement Learning, Bit by Bit
- A Bit Better? Quantifying Information for Bandit Learning
- Weak Signal Asymptotics for Sequentially Randomized Experiments
- Analysis and Design of Thompson Sampling for Stochastic Partial Monitoring
- Deciding What to Learn: A Rate-Distortion Approach
- Apple Tasting Revisited: Bayesian Approaches to Partially Monitored Online Binary Classification
- Best-of-Both-Worlds Algorithms for Partial Monitoring