machine learning

Kernel weighted importance sampling for off-policy evaluation in contextual bandits

arXiv:2607.15067

summary

The paper introduces Kernel-WIS, a new estimator that uses kernel-weighted importance sampling to evaluate policies offline in contextual bandit settings, offering consistency and better performance than existing methods, especially when the behavior policy is misspecified.

Abstract

This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including weighted importance sampling), particularly under behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of weighted importance sampling with the linearity of vanilla importance sampling.

Topics & keywords

#contextual bandits#off-policy evaluation#importance sampling#kernel methods#offline learningkernel weighted importance samplingasymptotic consistencybehavior policy misspecificationoff-policy evaluationcontextual bandits
Kernel weighted importance sampling for off-policy evaluation in contextual bandits · wovepaper