Optimal Best-Arm Identification Methods for Tail-Risk Measures
arXiv:2008.07606
Abstract
Conditional value-at-risk (CVaR) and value-at-risk (VaR) are popular tail-risk measures in finance and insurance industries as well as in highly reliable, safety-critical uncertain environments where often the underlying probability distributions are heavy-tailed. We use the multi-armed bandit best-arm identification framework and consider the problem of identifying the arm from amongst finitely many that has the smallest CVaR, VaR, or weighted sum of CVaR and mean. The latter captures the risk-return trade-off common in finance. Our main contribution is an optimal -correct algorithm that acts on general arms, including heavy-tailed distributions, and matches the lower bound on the expected number of samples needed, asymptotically (as approaches ). The algorithm requires solving a non-convex optimization problem in the space of probability measures, that requires delicate analysis. En-route, we develop new non-asymptotic empirical likelihood-based concentration inequalities for tail-risk measures which are tighter than those for popular truncation-based empirical estimators.
55 pages, 4 figures
References in corpus (8)
- Further Optimal Regret Bounds for Thompson Sampling
- Risk-Aversion in Multi-armed Bandits
- Sequential estimation of quantiles with applications to A/B-testing and best-arm identification
- Fairness risk measures
- Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewards
- Pure Exploration with Multiple Correct Answers
- Optimal -Correct Best-Arm Selection for Heavy-Tailed Distributions
- Regret Minimization in Heavy-Tailed Bandits