1 paper
Siyu Chen, Yitan Wang, Zhaoran Wang +1
We study the offline contextual bandit problem, where we aim to acquire an optimal policy using observational data. However, this data usually contains two deficiencies: (i) some v…