Functional Sequential Treatment Allocation with Covariates
arXiv:2001.10996 · doi:10.1017/S0266466623000051
Abstract
We consider a multi-armed bandit problem with covariates. Given a realization of the covariate vector, instead of targeting the treatment with highest conditional expectation, the decision maker targets the treatment which maximizes a general functional of the conditional potential outcome distribution, e.g., a conditional quantile, trimmed mean, or a socio-economic functional such as an inequality, welfare or poverty measure. We develop expected regret lower bounds for this problem, and construct a near minimax optimal assignment policy.
The material in this paper replaces the material on covariates in [v5] of "Functional Sequential Treatment Allocation"
References in corpus (9)
- Fast learning rates for plug-in classifiers
- Risk-Aversion in Multi-armed Bandits
- Risk-Averse Multi-Armed Bandit Problems under Mean-Variance Measure
- Nonparametric Bandits with Covariates
- Clinical trial design enabling ε-optimal treatment rules
- What Doubling Tricks Can and Can't Do for Multi-Armed Bandits
- Generalized Risk-Aversion in Stochastic Multi-Armed Bandits
- Optimal sequential treatment allocation
- Decision Variance in Online Learning