Exploration vs Exploitation vs Safety: Risk-averse Multi-Armed Bandits
arXiv:1401.1123
Abstract
Motivated by applications in energy management, this paper presents the Multi-Armed Risk-Aware Bandit (MARAB) algorithm. With the goal of limiting the exploration of risky arms, MARAB takes as arm quality its conditional value at risk. When the user-supplied risk level goes to 0, the arm quality tends toward the essential infimum of the arm distribution density, and MARAB tends toward the MIN multi-armed bandit algorithm, aimed at the arm with maximal minimal value. As a first contribution, this paper presents a theoretical analysis of the MIN algorithm under mild assumptions, establishing its robustness comparatively to UCB. The analysis is supported by extensive experimental validation of MIN and MARAB compared to UCB and state-of-art risk-aware MAB algorithms on artificial and real-world problems.
16 pages
References in corpus (5)
Cited by in corpus (22)
- Counterfactual Risk Minimization: Learning from Logged Bandit Feedback
- Risk-Averse Multi-Armed Bandit Problems under Mean-Variance Measure
- Generalized Risk-Aversion in Stochastic Multi-Armed Bandits
- Thompson Sampling Algorithms for Mean-Variance Bandits
- Learning Bounds for Risk-sensitive Learning
- X-Armed Bandits: Optimizing Quantiles, CVaR and Other Risks
- Learning for Dose Allocation in Adaptive Clinical Trials with Safety Constraints
- Risk-Constrained Thompson Sampling for CVaR Bandits
- Spectral risk-based learning using unbounded losses
- Learning to Map for Active Semantic Goal Navigation
- Learning Modular Safe Policies in the Bandit Setting with Application to Adaptive Clinical Trials
- The Fragility of Optimized Bandit Algorithms
- Old Dog Learns New Tricks: Randomized UCB for Bandit Problems
- Risk-Averse Action Selection Using Extreme Value Theory Estimates of the CVaR
- Decision Variance in Online Learning
- Generalized Batch Normalization: Towards Accelerating Deep Neural Networks
- Continuous Mean-Covariance Bandits
- Concentration bounds for CVaR estimation: The cases of light-tailed and heavy-tailed distributions
- Model-Free Risk-Sensitive Reinforcement Learning
- Risk averse non-stationary multi-armed bandits
- Multi-armed Bandits with Cost Subsidy
- Risk-Aware Algorithms for Combinatorial Semi-Bandits