1 paper · 1 filter
Zeyang Jia, Kosuke Imai, Michael Lingzhi Li
We introduce the cram method as a general statistical framework for evaluating the final learned policy from a multi-armed contextual bandit algorithm, using the dataset generated…