1 paper
Tong Li, Thiago de Queiroz Casanova, Eric M. Schwartz +3
Real-world contextual bandit problems with complex reward models are often tackled with iteratively trained models, such as boosting trees. However, it is difficult to directly app…