1 paper
Yuan Xie, Boyi Liu, Qiang Liu +3
When learning from a batch of logged bandit feedback, the discrepancy between the policy to be learned and the off-policy training data imposes statistical and computational challe…